> Markdown version of [/jobs/ext/622110-ml-infrastructure-mlops-engineer](https://www.wearedevelopers.com/jobs/ext/622110-ml-infrastructure-mlops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Infrastructure & MLOps Engineer - **Company:** AllSTEM Connections - **Location:** Ontario, CA, United States - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Airflow, Program Optimization, Continuous Delivery, Continuous Integration, Information Engineering, Distributed Systems, Azure Machine Learning, Cloud Platform System, Delivery Pipeline, Reliability of Systems, Backend, AI Platforms, Kubernetes, Deployment Automation, Machine Learning Operations - **Published:** June 24, 2026 - **Apply:** https://www.dice.com/job-detail/5d969548-f82c-4093-b8ab-962255785517 ## About the Role Experience: o5 to 10+ years of hands-on experience designing and operating large-scale distributed ML platforms. oProven track record of supporting production-grade ML workflows in cloud environments. Technical Mastery: oDeep expertise in container orchestration, specifically GKE (Google Kubernetes Engine) or equivalent enterprise Kubernetes environments. oHands-on experience building scalable ML pipelines (e.g., Kubeflow, Airflow, TFX). oStrong proficiency in distributed training strategies, feature store management, and model serving infrastructure. Soft Skills & Attributes: oPragmatic Mindset: Strong ownership-driven work style focused on consistency, system reliability, and cost-awareness. oEffective Communicator: Ability to collaborate seamlessly with highly technical researchers and platform engineers alike. Preferred Qualifications Prior experience working within dedicated, tier-1 enterprise ML/AI platform teams. Deep knowledge of distributed systems backend optimization and infrastructure-as-code (IaC). ## Description ML Infrastructure & Container Orchestration Distributed Clusters: Architect and maintain high-performance training and serving infrastructure utilizing Google Kubernetes Engine (GKE). Model Optimization: Design and implement high-efficiency optimization pipelines, including advanced knowledge distillation and foundational training tooling. Platform Scaling: Build, monitor, and optimize shared ML systems to ensure maximum infrastructure uptime, pipeline reliability, and cloud cost-efficiency. Data Engineering & Pipeline Automation Workflow Automation: Build robust, automated pipelines for standardized model training, validation, and continuous deployment (CI/CD for ML). Feature Platforms: Develop scalable data sampling and feature-generation platforms to accelerate research experimentation cycles. Onboarding & Usability: Drive high platform adoption by building intuitive, standardized deployment tools that decrease onboarding speed for research and engineering teams. Collaboration & Governance Cross-Functional Bridge: Collaborate closely with ML researchers and core software engineers to translate theoretical models into highly scalable production systems. Methodical Execution: Apply a disciplined, data-backed approach to identify infrastructure bottlenecks, reduce time-to-market, and stabilize complex deployments. ## Related Videos - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)