> Markdown version of [/jobs/ext/3003722-principal-ai-ml-platform-engineer](https://www.wearedevelopers.com/jobs/ext/3003722-principal-ai-ml-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal AI/ML Platform Engineer - **Company:** Unitedhealth Group Inc - **Location:** Washington, DC, United States (Remote available) - **Experience:** Experienced - **Salary:** $164,600.0 - $282,200.0 - **Contract:** Permanent contract - **Skills:** LTE (Telecommunication), Artificial Intelligence, Amazon Web Services, Computing Platforms, Microsoft Azure, DevOps, Distributed Computing Environment, Firmware, Infrastructure as a Service (IaaS), IBM Storage, Identity and Access Management, InfiniBand, Key Management, OAuth, Octopus Deploy, OpenID, OpenShift, Platform as a Service (PAAS), Role-Based Access Control, Remote Direct Memory Access, Red Hat Enterprise Linux, Reliability Engineering, Cloud Services, Azure Machine Learning, Weka, AI Infrastructure, Ceph (Software), Graphics Processing Unit (GPU), Google Cloud, Computer Network Technologies, Pytorch, Large Language Models, Multi-Agent Systems, AI Platforms, Kubernetes, Low Latency, Bare Metal, Free and Open-Source Software, TensorRT - **Published:** September 19, 2026 - **Apply:** https://www.dice.com/job-detail/808c29f1-e47f-479a-b8f8-cb4a853bba7d ## About the Role * Bachelor's degree or 4+ years of equivalent software/platform engineering experience in lieu of a degree * 10+ years of experience in infrastructure, DevOps, SRE, or ML platform engineering * 5+ years of experience operating Kubernetes or OpenShift at scale in production bare-metal or enterprise cloud environments * 3+ years of experience designing and managing accelerated-compute (AI/GPU) infrastructure utilizing NVIDIA GPU Operator, NFD, MIG/time-slicing, and DCGM * 3+ years of experience architecting hybrid AI platforms spanning self-hosted IaaS and public-cloud PaaS managed AI services (e.g., Azure AI Foundry, AWS Bedrock, or Google Cloud Platform Vertex AI) * 3+ years of experience with HPC/AI networking technologies, including InfiniBand or RoCEv2, GPUDirect RDMA, and NCCL collective communication limits * 3+ years of experience managing multi-cluster fleets using RHACM (or equivalent) and GitOps tooling (Argo CD or Flux) * 3+ years of experience implementing enterprise security and IAM controls for software workloads (RBAC, OIDC/OAuth, Vault secrets management, mTLS), * Experience with distributed training frameworks (PyTorch DDP/FSDP, DeepSpeed, Ray, JAX) and batch scheduling systems (Kueue, Volcano) * Hands-on experience with LLM inference serving technologies (vLLM, TensorRT-LLM, KServe) and platform tooling such as OpenShift AI (RHOAI) or Kubeflow pipelines * Experience with high-performance parallel storage systems (Ceph/ODF, Lustre, IBM Storage Scale, VAST, WEKA) * Active Red Hat certifications (e.g., Red Hat Certified Architect / RHCA) or open-source contributions to CNCF, OpenShift, or AI infrastructure projects * Experience operating AI/ML platforms within regulated healthcare environments under HIPAA and UHG data privacy controls *All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy. ## Description As a Principal AI/ML Platform Engineer on the UnitedHealth Group (UHG) enterprise team, you will serve as the AI Architect across IaaS and PaaS environments, owning the technical direction, reference architecture, and long-term evolution of our multi-tenant AI compute platform end to end. Our team builds and maintains an advanced compute estate spanning on-premises bare-metal Red Hat OpenShift AI clusters equipped with high-performance NVIDIA GPUs and InfiniBand/RoCE training fabrics, alongside public-cloud managed AI services including Azure AI Foundry, AWS Bedrock, and Google Cloud Platform Vertex AI. In this role, you will define architecture standards, optimize high-throughput model training and inference pipelines, establish cost and utilization economics, and enforce strict HIPAA, security, and data-governance standards for regulated healthcare workloads. You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week., * Own the end-to-end reference architecture for multi-tenant AI compute platforms across hybrid on-premises bare-metal OpenShift AI clusters and public-cloud managed AI platforms (Azure AI Foundry, AWS Bedrock, Google Cloud Platform Vertex AI) * Set network, latency, and topology standards for distributed training (NVLink, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL), ensuring interconnect boundaries are strictly maintained * Establish cluster governance, GitOps workflows (Argo CD), RHACM policies, and automated lifecycle management for bare-metal accelerated compute nodes * Design cost and utilization models including capex amortization, accelerator-sharing strategies (MIG/time-slicing), and cost-per-token/training-run math to inform accelerator procurement roadmaps * Standardize model-serving platforms (vLLM, KServe) and inference gateways, setting quantization policies, provenance review gates, and intelligent model routing rules * Define workload placement frameworks to determine self-hosted versus managed cloud deployment based on data residency, latency, cost, and compliance requirements * Design and enforce identity, access, and security controls for AI workloads and autonomous agents, including least-privilege RBAC, Vault secret management, short-lived credentials, and mTLS * Establish enterprise Service Level Objectives (SLOs), disaster recovery plans, and upgrade strategies for OpenShift, OpenShift AI, GPU operators, drivers, and firmware * Partner with cross-functional AI teams, LLM gateway engineers, privacy, and security stakeholders to ensure seamless integration and HIPAA compliance You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in. ## Related Videos - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) - [Delay the AI Overlords: How OAuth and OpenFGA Can Keep Your AI Agents from Going Rogue](https://www.wearedevelopers.com/videos/1637-delay-the-ai-overlords-how-oauth-and-openfga-can-keep-your-ai-agents-from-going-rogue) - [Delegating the chores of authenticating users to Keycloak](https://www.wearedevelopers.com/videos/1558-delegating-the-chores-of-authenticating-users-to-keycloak) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts)