> Markdown version of [/jobs/ext/223652-senior-azure-cloud-infrastructure-engineer-healthcare-ai-platform](https://www.wearedevelopers.com/jobs/ext/223652-senior-azure-cloud-infrastructure-engineer-healthcare-ai-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Azure Cloud Infrastructure Engineer (Healthcare AI Platform) - **Company:** Civie LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Audit Trail, Microsoft Azure, Backup Devices, Bash Shell, Cloud Computing, DevOps, Disaster Recovery, Distributed Computing Environment, Failover, Fault Tolerance, Python (Programming Language), Windows PowerShell, Role-Based Access Control, Azure Active Directory, Zero Trust Network Access, Azure Machine Learning, Virtual Machines, IBM Watson Health, Data Logging, Scripting, Cloud Platform System, System Availability, Parallel Computation, Containerization, Kubernetes, Bicep, Machine Learning Operations, Terraform, Virtual Private Clouds - **Published:** May 23, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=ae7fb2c0d94cfa42 ## About the Role Do you have experience in Virtual Private Clouds?, * 5-8+ years of hands-on experience with Microsoft Azure cloud infrastructure * Proven experience designing high-availability and disaster recovery systems in regulated environments * Strong background in healthcare or other compliance-heavy industries * Deep expertise in: * + Azure Virtual Machines, VM Scale Sets, and GPU compute * + Azure networking (VNets, Private Link, ExpressRoute, firewalls) * + Storage solutions (Blob, Files, managed disks with redundancy options) * Experience implementing compliance frameworks such as HIPAA or SOC 2 * Strong knowledge of identity and access control (RBAC, Azure AD, managed identities) * Experience with Kubernetes (AKS) and containerized workloads * Proficiency in scripting (Python, Bash, PowerShell), * Experience with Azure AI ecosystem (Azure Machine Learning, Azure AI Foundry, Cognitive Services) * Familiarity with distributed training, model parallelism, and GPU orchestration * Experience implementing MLOps pipelines in regulated environments * Azure certifications (Solutions Architect Expert, Security Engineer Associate, DevOps Engineer Expert) * Experience with zero-downtime deployments and blue/green or canary strategies Infrastructure Expectations * Multi-region architecture with automated failover * End-to-end encryption (data at rest and in transit) * Segmented environments (dev/staging/prod) with strict isolation * Real-time monitoring and alerting with defined SLAs * Automated backup and recovery with regular testing * Cost visibility and governance across all resources ## Description We're looking for a Senior Azure Cloud Infrastructure Engineer to design, build, and operate a highly resilient, secure, and cost-efficient cloud platform supporting advanced AI workloads in a healthcare environment. This role is responsible for mission-critical infrastructure powering our proprietary foundational AI model, including GPU-based compute, while meeting strict requirements for compliance, data protection, and high availability. You will play a key role in ensuring our systems are fault-tolerant, auditable, and continuously optimized for both performance and cost. What You'll Do * Architect and manage highly available, fault-tolerant systems on Microsoft Azure with multi-region redundancy and disaster recovery * Design infrastructure with strict adherence to healthcare compliance standards (e.g., HIPAA, HITRUST, SOC 2) * Provision and optimize GPU-based environments for AI/ML workloads, including large-scale model training and inference * Build secure, zero-trust architectures (private networking, encryption, identity isolation, least privilege access) * Implement backup, failover, and business continuity strategies with clearly defined RTO/RPO targets * Continuously reduce infrastructure costs through intelligent scaling, reserved capacity, spot instances, and workload optimization * Develop Infrastructure as Code (Terraform, Bicep, ARM) for repeatable, auditable deployments * Partner with AI/ML teams to productionize and scale foundational models reliably * Establish observability across systems (logging, monitoring, alerting) with proactive incident response * Conduct architecture reviews, risk assessments, and security audits, * Near-zero downtime systems with tested failover capabilities * Full compliance readiness with audit trails and documentation * Efficient GPU utilization supporting AI workloads at scale * Measurable reduction in cloud spend without compromising reliability or security * Seamless collaboration between infrastructure and AI teams Why This Role Matters You will be building the backbone of the next-generation healthcare AI platform - where reliability, security, and performance directly impact real-world outcomes. This is not just infrastructure; it is critical systems engineering at the intersection of cloud, AI and healthcare. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [System Resilience: Surviving the Software Storm](https://www.wearedevelopers.com/videos/874-system-resilience-surviving-the-software-storm) - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Building Sovereign AI: Lessons from Deploying Secure RAG Systems using Confidential Computing](https://www.wearedevelopers.com/videos/100108-building-sovereign-ai-lessons-from-deploying-secure-rag-systems-using-confidential-computing) - [Infrastructure as Prompts: Creating Azure Infrastructure with AI Agents](https://www.wearedevelopers.com/videos/1533-infrastructure-as-prompts-creating-azure-infrastructure-with-ai-agents) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)