Distributed Systems and ML Infrastructure Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
We’re looking for a Distributed Systems and ML Infrastructure Engineer to build the core services of a containerized, API-first AI platform. This is a role for someone who wants to build the thing itself, not integrate someone else’s.
You design and implement the services the platform runs on - workflow orchestration, data ingestion, results management, model serving, policy enforcement, usage accounting, audit logging. Those services have to hold up across cloud, dedicated, isolated, and limited-connectivity deployments, which means portability and operability are design constraints from the first commit rather than problems handed to someone downstream.
Development happens primarily on unrestricted infrastructure with an open-source toolchain. Engineers with the right access also carry releases into controlled production environments, integrate data sources there, and validate the platform in place - so there’s a path to seeing your work through to where it actually runs.
This position is contingent upon contract award. Travel of up to 15% may be required, primarily to Government facilities and between company locations. Unclassified work may be performed remotely, while classified promotion and validation activities require onsite work in an accredited facility and the appropriate security clearance., * Design and implement platform services for workflow orchestration, data ingest, and results management, exposed through documented APIs with no proprietary front end
- Implement and operate model gateway and serving services that route invocations to approved managed model services with policy enforcement, usage accounting, and audit logging
- Maintain a documented provider abstraction so the platform runs on AWS-native managed services where appropriate while remaining deployable across other cloud and dedicated environments
- Deploy platform releases into classified host environments, perform data source integration, and execute validation procedures on a recurring promotion cadence
- Verify environment parity after each promotion
- Size and validate the platform against documented workload models, and verify capacity and performance by load test
- Constrain platform dependencies to services confirmed available in the target environments, and gate any development-only dependency behind feature flags
- Reproduce high-side defects on the low side through sanitized feedback paths and fix them where the full toolchain is available, Oversee DHMO quality management and utilization management governance, regulatory committees, audit readiness, corrective action monitoring, compliance reporting, and quality improvement. Analyze operational, financial, regulatory, satisfaction, grievance, and performance data to identify risks and trends. Develop governance processes, policies, reporting frameworks, and documentation standards. Partner with cross-functional stakeholders and leadership to improve compliance, organizational performance, member experience, and provider outcomes., Manages strategic relationships with national commercial fleet clients, focusing on growth, retention, operational performance, and customer satisfaction. Monitors fleet KPIs, analyzes data, prepares reports, leads business reviews, resolves escalations, and coordinates cross-functional teams. Identifies process improvements, oversees SLA and audit performance, supports programs and SOPs, and provides proactive communication regarding operations, service issues, and product enhancements.
Requirements
Clearance: An active TS/SCI clearance with CI polygraph is strongly preferred. Candidates without an active clearance may be considered for unclassified work but must be eligible to obtain and maintain a U.S. security clearance., * U.S. citizenship and eligibility to obtain and maintain a U.S. security clearance
- 6+ years of experience in distributed systems, platform engineering, infrastructure engineering, or a related software engineering role
- Production experience operating Kubernetes and containerized workloads on a major cloud platform
- Experience with managed Kubernetes services such as Amazon EKS or an equivalent platform
- Experience supporting machine learning workloads in production, such as model serving, GPU scheduling, or large-scale data and evaluation pipelines
- Experience designing, building, or operating distributed services that support reliability, scalability, and performance requirements
- Proficiency in Python, Go, or a comparable programming language
- Experience with infrastructure-as-code and deployment tools such as Terraform, Helm, or equivalent technologies
- Experience designing API-first services and implementing documented interface specifications
- Experience testing platform capacity and performance against expected workload requirements
- Ability to document technical interfaces, deployment procedures, architectural decisions, and validation results
- Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience
Nice to Have
- Active TS/SCI clearance with CI polygraph
- Hands-on experience deploying or operating software in classified, air-gapped, or limited-connectivity environments
- Experience with classified cloud environments, including AWS Secret or Top Secret regions
- Experience building or operating model gateways, LLM routing layers, inference services, or inference-brokering platforms
- Familiarity with software promotion into classified environments, cross-domain transfer processes, and release packaging
- Experience with agentic workflow frameworks or the orchestration of multi-step AI pipelines
- Experience supporting GPU-accelerated workloads and distributed model inference
- Experience designing vendor-agnostic platforms that operate across multiple cloud or dedicated environments
- Experience supporting rapid prototyping programs or defense innovation initiatives
Benefits & conditions
In-Office 2 Locations 145K-250K Annually Senior level In-Office 2 Locations 145K-250K Annually Senior level Build and operate containerized, API-first AI platform services for workflow orchestration, data ingestion, model serving, policy enforcement, audit logging, and usage accounting. Deploy and validate releases across cloud, dedicated, classified, air-gapped, and limited-connectivity environments. Design for portability, reliability, scalability, and performance; conduct load testing, verify environment parity, integrate data sources, and resolve high-side defects through sanitized feedback paths., What We Offer
- Medical, Dental & Vision - 100% paid for employees, 75% for dependents
- 401(k) Match - Up to 5% with full vesting after 2 years
- Unlimited PTO - With a required minimum of 15 days off annually
- Fully Remote Setup - Includes up to $3,000 equipment reimbursement
- Continuous Education - Includes up to $500 reimbursement
- Disability & Life Insurance - 100% employer-paid
- HSA & FSA Options - With monthly HSA contributions from OpenTeams
Grow With Us
At OpenTeams, growth isn’t just about the company-it’s about you. We believe the best careers are built at the edge of your potential. That is where new tools, ideas, and technologies change the world. Here, you’ll work alongside pioneers of AI, solving problems that matter: making AI more transparent, more ethical, and more empowering. As your skills grow, our career framework provides a pathway and recognition of that increased impact.
Opportunities aren’t limited by geography. You’ll collaborate with global experts, contribute to open source projects that power the world’s technology, and stretch your skills daily. That global perspective and diversity makes our solution more universal and robust. We are committed to continuing to celebrate diversity on our team.
Supported people are successful people. We offer 100% employer paid medical premiums for employees and self-managed PTO with a minimum time off requirement, so that our teams are able to do their best work. We invest in curiosity, creativity, and ownership. That means you’ll be trusted to boldly innovate, supported to learn fast, and celebrated for successful collaboration. Commitment to diversity, equity, inclusion, and belonging
OpenTeams understands that valuing diverse creative practices and forms of knowledge is crucial to and enriches the company’s core mission. We encourage applications from everyone, including members of all equity-seeking communities, such as (but certainly not limited to) women, racialized and Indigenous persons, disabled people, persons of all sexual orientations, gender identities and expressions., In-Office Colorado Springs, CO, USA 112K-185K Annually Senior level 112K-185K Annually Senior level Aerospace * Information Technology * Software * Cybersecurity * Design * Defense * Manufacturing Designs and evaluates complex aerospace and space mission systems from a human performance perspective. Defines human engineering requirements, develops research and analysis plans, evaluates prototypes, conducts user interviews and field studies, and integrates human factors across multidisciplinary teams. Applies military standards, human performance principles, workload and situation-awareness assessments, task analysis, workflow mapping, and usability methods. The role is fully onsite in Colorado Springs, requires occasional travel and after-hours support, and requires eligibility for U.S. Top Secret/SCI clearance and special program access. Top Skills: Agile DevelopmentAstro Space Ux Design SystemDodi 5000.95Human Performance ModelingHuman System Integration (Hsi)Jssg-2010Mil-Std-1472Mil-Std-1474Mil-Std-1787Mil-Std-2525Mil-Std-3009Mil-Std-411FModel-Based Systems Engineering (Mbse)Nasa TlxUsability SoftwareUser Interface Prototyping Tools
What you need to know about the Colorado Tech Scene
With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.
Key Facts About Colorado Tech
- Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
- Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
- Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
- Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute
About the company
Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it. OpenTeams exists to make ownership possible. Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves. If that sounds like your kind of work, we’d like to meet you.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
MLOps – What’s the deal behind it?
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
Highest Paying Tech Companies for Developers