Senior Azure Cloud Infrastructure Engineer (Healthcare AI Platform)

Civie LLC
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Systems Engineering Audit Trail Microsoft Azure Backup Devices Bash Shell Cloud Computing DevOps Disaster Recovery Distributed Computing Environment Failover Fault Tolerance
+19 more
Python (Programming Language) Windows PowerShell Role-Based Access Control Azure Active Directory Zero Trust Network Access Azure Machine Learning Virtual Machines IBM Watson Health Data Logging Scripting Cloud Platform System System Availability Parallel Computation Containerization Kubernetes Bicep Machine Learning Operations Terraform Virtual Private Clouds

Job description

We’re looking for a Senior Azure Cloud Infrastructure Engineer to design, build, and operate a highly resilient, secure, and cost-efficient cloud platform supporting advanced AI workloads in a healthcare environment.

This role is responsible for mission-critical infrastructure powering our proprietary foundational AI model, including GPU-based compute, while meeting strict requirements for compliance, data protection, and high availability. You will play a key role in ensuring our systems are fault-tolerant, auditable, and continuously optimized for both performance and cost.

What You’ll Do

  • Architect and manage highly available, fault-tolerant systems on Microsoft Azure with multi-region redundancy and disaster recovery
  • Design infrastructure with strict adherence to healthcare compliance standards (e.g., HIPAA, HITRUST, SOC 2)
  • Provision and optimize GPU-based environments for AI/ML workloads, including large-scale model training and inference
  • Build secure, zero-trust architectures (private networking, encryption, identity isolation, least privilege access)
  • Implement backup, failover, and business continuity strategies with clearly defined RTO/RPO targets
  • Continuously reduce infrastructure costs through intelligent scaling, reserved capacity, spot instances, and workload optimization
  • Develop Infrastructure as Code (Terraform, Bicep, ARM) for repeatable, auditable deployments
  • Partner with AI/ML teams to productionize and scale foundational models reliably
  • Establish observability across systems (logging, monitoring, alerting) with proactive incident response
  • Conduct architecture reviews, risk assessments, and security audits, * Near-zero downtime systems with tested failover capabilities
  • Full compliance readiness with audit trails and documentation
  • Efficient GPU utilization supporting AI workloads at scale
  • Measurable reduction in cloud spend without compromising reliability or security
  • Seamless collaboration between infrastructure and AI teams

Why This Role Matters

You will be building the backbone of the next-generation healthcare AI platform - where reliability, security, and performance directly impact real-world outcomes. This is not just infrastructure; it is critical systems engineering at the intersection of cloud, AI and healthcare.

Requirements

Do you have experience in Virtual Private Clouds?, * 5-8+ years of hands-on experience with Microsoft Azure cloud infrastructure

  • Proven experience designing high-availability and disaster recovery systems in regulated environments
  • Strong background in healthcare or other compliance-heavy industries
  • Deep expertise in: *

  • Azure Virtual Machines, VM Scale Sets, and GPU compute *

  • Azure networking (VNets, Private Link, ExpressRoute, firewalls) *

  • Storage solutions (Blob, Files, managed disks with redundancy options)
  • Experience implementing compliance frameworks such as HIPAA or SOC 2
  • Strong knowledge of identity and access control (RBAC, Azure AD, managed identities)
  • Experience with Kubernetes (AKS) and containerized workloads
  • Proficiency in scripting (Python, Bash, PowerShell), * Experience with Azure AI ecosystem (Azure Machine Learning, Azure AI Foundry, Cognitive Services)
  • Familiarity with distributed training, model parallelism, and GPU orchestration
  • Experience implementing MLOps pipelines in regulated environments
  • Azure certifications (Solutions Architect Expert, Security Engineer Associate, DevOps Engineer Expert)
  • Experience with zero-downtime deployments and blue/green or canary strategies

Infrastructure Expectations

  • Multi-region architecture with automated failover
  • End-to-end encryption (data at rest and in transit)
  • Segmented environments (dev/staging/prod) with strict isolation
  • Real-time monitoring and alerting with defined SLAs
  • Automated backup and recovery with regular testing
  • Cost visibility and governance across all resources

Benefits & conditions

Pulled from the full job description

  • 401(k) matching
  • Vision insurance
  • Health savings account
  • Dental insurance
  • Life insurance
  • Disability insurance
  • Profit sharing, * Paid vacation, sick time, and personal days
  • 11 company paid holidays
  • Quarterly UberEats voucher
  • Monthly Fringe benefits
  • Flexible work schedules
  • Professional development stipend
  • Health, dental, and vision benefits, with employer HSA contribution
  • STD, LTD and life insurance
  • 401(k) company match and profit sharing

About the company

CIVIE was founded in 2018 to tackle the toughest radiology challenges with advanced and evolving technologies. Designed by a team with deep radiology industry experience, CIVIE is driven by our passion for innovation.

You’’ be joining a fully remote team of experts who are deeply committed to our mission and focused on delivering solutions that truly move the needle for our partners and the healthcare industry. If what we are doing excites you, we would love to connect.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:59 min

Scaling clusters and handling automated replica failover

Jürgen Pilz · WWC 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

5:25 min

Implementing redundancy, failover, and architectural load balancing patterns

Mihaela-Roxana Ghidersa · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all