Infrastructure Cloud Engineer (Azure/AWS)

CAI, Inc.
Columbia, SC, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$110,000.0 - $135,000.0
Working hours
Regular working hours

Tech stack

.NET Framework Amazon Web Services Application Performance Management Microsoft Azure Backup Devices Cloud Computing Cloud Engineering Data as a Services Relational Databases DevOps Key Management PostgreSQL
+15 more
Log Analysis SQL Azure Windows Servers Virtual Desktops Network Configuration and Change Management Performance Tuning Redis Reliability Engineering Datadog Cloud Monitoring Grafana Caching Deployment Automation Bicep Terraform

Job description

We are looking for an Infrastructure Cloud Engineer to own the reliability, performance, and day-to-day operation of the Azure and AWS hosted services that run our internal and client-facing applications. This is a hands-on, cloud-native role. You will spend your time running managed Azure and AWS services well rather than building infrastructure platforms from the ground up. Think pragmatic reliability engineering for a lean, focused IT organization: you keep the systems healthy, make them observable, automate the toil out of them, and get pulled in when something breaks.

You will partner closely with application developers, solution architects, and the broader infrastructure team. If you like owning production, care about clean observability, and want to work across a real portfolio of services without the overhead of a giant platform org, this is a good fit. This position will be full-time and remote.

“This position does not offer employment sponsorship. All candidates must be eligible to work without need for sponsorship by employer.”

What You’ll Do

  • Operate and maintain our Azure Container Apps (ACA) environments: deployments, scaling rules, revisions, ingress, networking, and secrets management
  • Own our managed data services in Azure, primarily Azure Database for PostgreSQL (Flexible Server) and Azure Cache for Redis: sizing, configuration, backups, patching, connection management, and performance tuning
  • Build and maintain observability across the stack: metrics, logs, traces, dashboards, and alerting using tools like Azure Monitor, Application Insights, Log Analytics, and Grafana
  • Define and track service health signals (availability, latency, error rates) and drive down noise so alerts mean something
  • Automate provisioning and configuration through infrastructure-as-code (Bicep or Terraform) and CI/CD pipelines
  • Participate in incident response: triage, mitigate, and lead or contribute to blameless postmortems, then close the loop with real fixes
  • Manage identity, access, and network configuration for hosted services, including Entra ID integration, private networking, and least-privilege access
  • Handle cost visibility and optimization for the services you own, right-sizing resources and flagging waste
  • Write and maintain runbooks and documentation so operational knowledge does not live in one person’s head
  • Collaborate with developers on deployment patterns, readiness for production, and reliability improvements to the applications running on your infrastructure

Requirements

  • 4 or more years in an infrastructure, DevOps, cloud operations, or SRE-adjacent role
  • Solid hands-on Azure and AWS experience, especially with container-based compute and managed data services
  • Practical experience running a relational database in production (PostgreSQL preferred) and understanding of caching with Redis
  • Real experience with observability tooling: you have built dashboards, set up meaningful alerts, and used telemetry to actually find and fix problems
  • Comfort with infrastructure-as-code and CI/CD pipelines
  • Scripting and automation skills, and a bias toward eliminating repetitive manual work
  • Sound judgment during incidents and a calm, methodical approach to troubleshooting under pressure
  • Clear written communication and a habit of documenting your work, * Azure certifications (AZ-104, AZ-400, or similar)
  • Exposure to compliance-driven environments (FedRAMP, StateRAMP, SOC 2, or public-sector requirements)
  • Experience supporting Windows Server
  • Experience supporting Azure Virtual Desktop
  • Experience supporting client-facing or externally exposed applications
  • Familiarity with .NET application hosting and deployment patterns
  • Prior participation in an on-call rotation

Physical Demands

  • Ability to safely and successfully perform the essential job functions
  • Sedentary work that involves sitting or remaining stationary most of the time with occasional need to move around the office to attend meetings, etc.
  • Ability to conduct repetitive tasks on a computer, utilizing a mouse, keyboard, and monitor

Benefits & conditions

$110,000 - $135,000 per year

The pay range for this position is listed above. Exact compensation may vary based on several factors, including location, experience, and education. Benefit packages include medical, dental, and vision insurance, as well as 401k retirement account access. Employees in this role receive paid time off and may also be entitled to paid sick leave and/or other paid time off as provided by applicable law.

About the company

CAI is a global services firm with over 9,000 associates worldwide and a yearly revenue of $1.3 billion+. We have over 40 years of excellence in uniting talent and technology to power the possible for our clients, colleagues, and communities. As a privately held company, we have the freedom and focus to do what is right-whatever it takes. Our tailor-made solutions create lasting results across the public and commercial sectors, and we are trailblazers in bringing neurodiversity to the enterprise.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · WWC 2022

Videos

See all

Related articles

See all