Infrastructure Engineer - AI/ML Platform

Openteams, Inc.
California, NC, United States
2 days ago
Apply on job-boards.greenhouse.io
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$145,000.0 - $250,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Systems Engineering Microsoft Azure CompTIA Security+ Computer Programming Continuous Integration DevOps Monitoring of Systems Identity and Access Management Python (Programming Language) Open Source Technology Program Analysis
+17 more
Reliability Engineering Prometheus Runbook Data Logging Pulumi Google Cloud System Availability Grafana Kubernetes Infrastructure Automation Frameworks Deployment Automation Machine Learning Operations Asynchronous Programming Terraform Data Pipelines Static Application Security Testing Dynamic Application Security Testing

Job description

We’re seeking a Senior Infrastructure Engineer to build and operate the platform underlying secure AI test and evaluation capabilities for Government teams assessing AI systems.

You’ll own a Kubernetes-based platform supporting demanding AI/ML workloads, including GPU scheduling, large-scale data movement, reproducible test execution, and multi-tenant isolation. The platform must operate reliably within Government environments that may have restricted networks, accreditation boundaries, limited connectivity, and no assumption of outbound internet access or managed cloud services.

You’ll use open-source technologies such as OpenTofu, Terraform, Helm, Argo CD, Kubernetes operators, and Nebari to build reusable and composable infrastructure. You’ll also own platform reliability, observability, capacity planning, upgrade paths, hardened configurations, and the documentation needed to deploy and maintain the platform.

This is a fully remote, U.S.-based role working with a distributed team that relies heavily on asynchronous communication., * Build and operate the Kubernetes platform supporting AI test and evaluation frameworks

  • Implement GPU scheduling, workload orchestration, resource management, and multi-tenant isolation across evaluation teams
  • Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines that make the platform reproducible from source
  • Develop reusable and modular infrastructure components that can be composed into independently owned and operated platforms
  • Contribute to Nebari and other open-source Kubernetes, infrastructure, and MLOps projects used by the platform
  • Own platform reliability, including capacity planning, upgrade strategies, failure-mode analysis, backup and recovery considerations, and operational readiness
  • Design and implement observability, monitoring, logging, tracing, and alerting for large-scale AI/ML workloads
  • Develop operational runbooks and documentation that enable other engineers to deploy, operate, and troubleshoot the platform
  • Deploy, configure, and harden infrastructure within secure, restricted, disconnected, or limited-connectivity Government environments
  • Support security authorization and compliance activities through infrastructure documentation, hardened configurations, control evidence, and repeatable deployment processes
  • Integrate automated security tooling for container scanning, static and dynamic analysis, artifact signing, and policy enforcement
  • Collaborate with Government stakeholders, security personnel, software engineers, and ML engineers to translate platform requirements into reliable infrastructure
  • Provide technical leadership, contribute to engineering standards, and mentor less-experienced team members
  • Collaborate effectively within a remote and distributed team using asynchronous communication practices, Design, build, and operate large-scale on-prem Kubernetes platforms (OpenShift/Anthos) for AI/ML and GPU workloads. Own control plane and etcd lifecycle, build controllers/operators in Go, automate via IaC, improve observability and resource utilization, enable ML training/inference/LLM deployments, participate in on-call incident response, and mentor engineers. Top Skills: AnthosCrdEtcdGoGpuKubeflowKubernetesKubernetes ControllersMlflowOpenshiftOperatorsWebhooks Liberty Mutual Insurance

Requirements

  • U.S. citizenship and ability to obtain and maintain a Secret security clearance
  • 6+ years of hands-on infrastructure, platform, DevOps, or site reliability engineering experience supporting production systems
  • Strong understanding of infrastructure engineering principles, including scalability, reliability, observability, security, and automation
  • Production experience with Kubernetes, including workload scheduling, resource management, and multi-tenant environments
  • Experience with automated security tooling, such as container scanning, SAST/DAST, artifact signing, and policy enforcement
  • Proficiency with infrastructure-as-code tools such as Terraform, OpenTofu, Pulumi, or equivalent technologies
  • Experience with at least one major cloud platform-AWS, Azure, or Google Cloud-including networking, security, storage, and compute services
  • Experience implementing monitoring and observability using tools such as OpenTelemetry, Prometheus, Grafana, or equivalent technologies
  • Strong programming or automation skills using Python, Go, or a comparable language
  • Experience with CI/CD practices, GitOps workflows, and infrastructure automation
  • Experience creating maintainable operational documentation, deployment procedures, and runbooks
  • Experience leading technical initiatives, establishing engineering practices, or mentoring other engineers
  • Ability to work independently and collaborate effectively within a remote, distributed team
  • Ability to give and receive constructive technical feedback

Nice to Have

  • Experience deploying or operating infrastructure in air-gapped, disconnected, or highly restricted environments
  • Experience engineering systems subject to the Risk Management Framework, NIST 800-53, NIST 800-171, or comparable security requirements
  • Experience supporting security personnel in achieving an Authorization to Operate for complex systems in Government or Intelligence Community environments
  • Experience supporting systems operating at Impact Level 5 or higher
  • Current DoD 8140 qualifying certification, such as Security+, CISSP, CISM, or an equivalent IAT/IAM Level II or III credential
  • Experience building MLOps pipelines or infrastructure supporting AI/ML workloads
  • Experience with GPU scheduling, large-scale data pipelines, or reproducible ML evaluation workloads
  • Experience with model-serving frameworks such as KServe, vLLM, LLM-D, or equivalent technologies
  • Familiarity with data sovereignty, privacy, and security requirements for enterprise or Government AI systems
  • Contributions to open-source Kubernetes, infrastructure, MLOps, or observability projects
  • Experience with Nebari

Benefits & conditions

Build and operate a secure Kubernetes-based AI/ML platform for government test and evaluation teams. Responsibilities include GPU scheduling, workload orchestration, infrastructure-as-code, GitOps, observability, capacity planning, upgrades, security hardening, compliance support, and operational documentation. The role involves restricted and disconnected environments, automated security tooling, collaboration with government and engineering stakeholders, technical leadership, and mentoring. The summary above was generated by AI, What We Offer

  • Medical, Dental & Vision - 100% paid for employees, 75% for dependents
  • 401(k) Match - Up to 5% with full vesting after 2 years
  • Unlimited PTO - With a required minimum of 15 days off annually
  • Fully Remote Setup - Includes up to $3,000 equipment reimbursement
  • Continuous Education - Includes up to $500 reimbursement
  • Disability & Life Insurance - 100% employer-paid
  • HSA & FSA Options - With monthly HSA contributions from OpenTeams

Grow With Us

At OpenTeams, growth isn’t just about the company-it’s about you. We believe the best careers are built at the edge of your potential. That is where new tools, ideas, and technologies change the world. Here, you’ll work alongside pioneers of AI, solving problems that matter: making AI more transparent, more ethical, and more empowering. As your skills grow, our career framework provides a pathway and recognition of that increased impact.

Opportunities aren’t limited by geography. You’ll collaborate with global experts, contribute to open source projects that power the world’s technology, and stretch your skills daily. That global perspective and diversity makes our solution more universal and robust. We are committed to continuing to celebrate diversity on our team.

Supported people are successful people. We offer 100% employer paid medical premiums for employees and self-managed PTO with a minimum time off requirement, so that our teams are able to do their best work. We invest in curiosity, creativity, and ownership. That means you’ll be trusted to boldly innovate, supported to learn fast, and celebrated for successful collaboration. Commitment to diversity, equity, inclusion, and belonging

OpenTeams understands that valuing diverse creative practices and forms of knowledge is crucial to and enriches the company’s core mission. We encourage applications from everyone, including members of all equity-seeking communities, such as (but certainly not limited to) women, racialized and Indigenous persons, disabled people, persons of all sexual orientations, gender identities and expressions., An Hour Ago Remote or Hybrid 45K-85K Annually Junior 45K-85K Annually Junior Artificial Intelligence * Fintech * Insurance * Marketing Tech * Software * Analytics Handle inbound calls and warm leads, consult customers on insurance needs, recommend appropriate property and casualty coverage, and convert prospects into policyholders. The role includes paid training and licensing, customer communication, sales closing, and brand representation. Representatives must work a fixed schedule including one weekend day and meet remote-work requirements with a dedicated workspace and wired high-speed internet. Top Skills: PcWired High-Speed Internet SOPHiA GENETICS, Own and expand customer relationships across the Western United States for a cloud-based healthcare genomics platform. Drive adoption, usage, and account growth within hospitals, laboratories, and healthcare organizations; identify expansion opportunities; navigate complex stakeholder groups; provide account support; and partner with sales teams on customer meetings and commercial development. The field-based role requires approximately 30-40% territory travel. Top Skills: Artificial IntelligenceCloud-Native PlatformsSophia Ddm Platform

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

About the company

Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it. OpenTeams exists to make ownership possible. Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves. If that sounds like your kind of work, we’d like to meet you.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on job-boards.greenhouse.io
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all