> Markdown version of [/jobs/ext/2695485-infrastructure-engineer-ai-ml-platform](https://www.wearedevelopers.com/jobs/ext/2695485-infrastructure-engineer-ai-ml-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Engineer - AI/ML Platform - **Company:** Openteams, Inc. - **Location:** California, NC, United States - **Experience:** Expert - **Salary:** $145,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Microsoft Azure, CompTIA Security+, Computer Programming, Continuous Integration, DevOps, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Open Source Technology, Program Analysis, Reliability Engineering, Prometheus, Runbook, Data Logging, Pulumi, Google Cloud, System Availability, Grafana, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Machine Learning Operations, Asynchronous Programming, Terraform, Data Pipelines, Static Application Security Testing, Dynamic Application Security Testing - **Published:** September 3, 2026 - **Apply:** https://job-boards.greenhouse.io/openteams/jobs/4729695005 ## About the Role * U.S. citizenship and ability to obtain and maintain a Secret security clearance * 6+ years of hands-on infrastructure, platform, DevOps, or site reliability engineering experience supporting production systems * Strong understanding of infrastructure engineering principles, including scalability, reliability, observability, security, and automation * Production experience with Kubernetes, including workload scheduling, resource management, and multi-tenant environments * Experience with automated security tooling, such as container scanning, SAST/DAST, artifact signing, and policy enforcement * Proficiency with infrastructure-as-code tools such as Terraform, OpenTofu, Pulumi, or equivalent technologies * Experience with at least one major cloud platform-AWS, Azure, or Google Cloud-including networking, security, storage, and compute services * Experience implementing monitoring and observability using tools such as OpenTelemetry, Prometheus, Grafana, or equivalent technologies * Strong programming or automation skills using Python, Go, or a comparable language * Experience with CI/CD practices, GitOps workflows, and infrastructure automation * Experience creating maintainable operational documentation, deployment procedures, and runbooks * Experience leading technical initiatives, establishing engineering practices, or mentoring other engineers * Ability to work independently and collaborate effectively within a remote, distributed team * Ability to give and receive constructive technical feedback Nice to Have * Experience deploying or operating infrastructure in air-gapped, disconnected, or highly restricted environments * Experience engineering systems subject to the Risk Management Framework, NIST 800-53, NIST 800-171, or comparable security requirements * Experience supporting security personnel in achieving an Authorization to Operate for complex systems in Government or Intelligence Community environments * Experience supporting systems operating at Impact Level 5 or higher * Current DoD 8140 qualifying certification, such as Security+, CISSP, CISM, or an equivalent IAT/IAM Level II or III credential * Experience building MLOps pipelines or infrastructure supporting AI/ML workloads * Experience with GPU scheduling, large-scale data pipelines, or reproducible ML evaluation workloads * Experience with model-serving frameworks such as KServe, vLLM, LLM-D, or equivalent technologies * Familiarity with data sovereignty, privacy, and security requirements for enterprise or Government AI systems * Contributions to open-source Kubernetes, infrastructure, MLOps, or observability projects * Experience with Nebari ## Description We're seeking a Senior Infrastructure Engineer to build and operate the platform underlying secure AI test and evaluation capabilities for Government teams assessing AI systems. You'll own a Kubernetes-based platform supporting demanding AI/ML workloads, including GPU scheduling, large-scale data movement, reproducible test execution, and multi-tenant isolation. The platform must operate reliably within Government environments that may have restricted networks, accreditation boundaries, limited connectivity, and no assumption of outbound internet access or managed cloud services. You'll use open-source technologies such as OpenTofu, Terraform, Helm, Argo CD, Kubernetes operators, and Nebari to build reusable and composable infrastructure. You'll also own platform reliability, observability, capacity planning, upgrade paths, hardened configurations, and the documentation needed to deploy and maintain the platform. This is a fully remote, U.S.-based role working with a distributed team that relies heavily on asynchronous communication., * Build and operate the Kubernetes platform supporting AI test and evaluation frameworks * Implement GPU scheduling, workload orchestration, resource management, and multi-tenant isolation across evaluation teams * Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines that make the platform reproducible from source * Develop reusable and modular infrastructure components that can be composed into independently owned and operated platforms * Contribute to Nebari and other open-source Kubernetes, infrastructure, and MLOps projects used by the platform * Own platform reliability, including capacity planning, upgrade strategies, failure-mode analysis, backup and recovery considerations, and operational readiness * Design and implement observability, monitoring, logging, tracing, and alerting for large-scale AI/ML workloads * Develop operational runbooks and documentation that enable other engineers to deploy, operate, and troubleshoot the platform * Deploy, configure, and harden infrastructure within secure, restricted, disconnected, or limited-connectivity Government environments * Support security authorization and compliance activities through infrastructure documentation, hardened configurations, control evidence, and repeatable deployment processes * Integrate automated security tooling for container scanning, static and dynamic analysis, artifact signing, and policy enforcement * Collaborate with Government stakeholders, security personnel, software engineers, and ML engineers to translate platform requirements into reliable infrastructure * Provide technical leadership, contribute to engineering standards, and mentor less-experienced team members * Collaborate effectively within a remote and distributed team using asynchronous communication practices, Design, build, and operate large-scale on-prem Kubernetes platforms (OpenShift/Anthos) for AI/ML and GPU workloads. Own control plane and etcd lifecycle, build controllers/operators in Go, automate via IaC, improve observability and resource utilization, enable ML training/inference/LLM deployments, participate in on-call incident response, and mentor engineers. Top Skills: AnthosCrdEtcdGoGpuKubeflowKubernetesKubernetes ControllersMlflowOpenshiftOperatorsWebhooks Liberty Mutual Insurance ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)