ML Infrastructure & MLOps Engineer

AllSTEM Connections
Ontario, CA, United States
about 2 months ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Airflow Program Optimization Continuous Delivery Continuous Integration Information Engineering Distributed Systems Azure Machine Learning Cloud Platform System Delivery Pipeline Reliability of Systems Backend AI Platforms
+3 more
Kubernetes Deployment Automation Machine Learning Operations

Job description

ML Infrastructure & Container Orchestration Distributed Clusters: Architect and maintain high-performance training and serving infrastructure utilizing Google Kubernetes Engine (GKE). Model Optimization: Design and implement high-efficiency optimization pipelines, including advanced knowledge distillation and foundational training tooling. Platform Scaling: Build, monitor, and optimize shared ML systems to ensure maximum infrastructure uptime, pipeline reliability, and cloud cost-efficiency.

Data Engineering & Pipeline Automation Workflow Automation: Build robust, automated pipelines for standardized model training, validation, and continuous deployment (CI/CD for ML). Feature Platforms: Develop scalable data sampling and feature-generation platforms to accelerate research experimentation cycles. Onboarding & Usability: Drive high platform adoption by building intuitive, standardized deployment tools that decrease onboarding speed for research and engineering teams.

Collaboration & Governance Cross-Functional Bridge: Collaborate closely with ML researchers and core software engineers to translate theoretical models into highly scalable production systems. Methodical Execution: Apply a disciplined, data-backed approach to identify infrastructure bottlenecks, reduce time-to-market, and stabilize complex deployments.

Requirements

Experience: o5 to 10+ years of hands-on experience designing and operating large-scale distributed ML platforms. oProven track record of supporting production-grade ML workflows in cloud environments. Technical Mastery: oDeep expertise in container orchestration, specifically GKE (Google Kubernetes Engine) or equivalent enterprise Kubernetes environments. oHands-on experience building scalable ML pipelines (e.g., Kubeflow, Airflow, TFX). oStrong proficiency in distributed training strategies, feature store management, and model serving infrastructure. Soft Skills & Attributes: oPragmatic Mindset: Strong ownership-driven work style focused on consistency, system reliability, and cost-awareness. oEffective Communicator: Ability to collaborate seamlessly with highly technical researchers and platform engineers alike.

Preferred Qualifications Prior experience working within dedicated, tier-1 enterprise ML/AI platform teams. Deep knowledge of distributed systems backend optimization and infrastructure-as-code (IaC).

About the company

For temporary assignments lasting 13 weeks or longer, AllSTEM Connections is pleased to offer major medical, dental, vision, 401k and any statutory sick pay where required.

We are committed to working with and providing reasonable accommodations to individuals with disabilities. If you need a reasonable accommodation for any part of the employment process, please contact your staffing representative who will reach out to our HR team.

AllSTEM Connections participates in the E-Verify program in certain locations as required by law. Learn more about the E-Verify program. _Participation_Poster_ES.pdf

We also consider for employment qualified applicants regardless of criminal histories, consistent with legal requirements, including, if applicable, the City of Los Angeles’ Fair Chance Initiative for Hiring Ordinance. Pursuant to applicable state and municipal Fair Chance Laws and Ordinances, we will consider for employment-qualified applicants with arrest and conviction records, including, if applicable, the San Francisco Fair Chance Ordinance. For Los Angeles, CA applicants: Qualified applications with arrest or conviction records will be considered for employment in accordance with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · WWC 2023

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · WWC 2024

2:44 min

Defining core roles and responsibilities in MLOps teams

Bas Geerdink · LIVE

Videos

See all

Related articles

See all