> Markdown version of [/jobs/ext/626977-machine-learning-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/626977-machine-learning-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Infrastructure Engineer - **Company:** AllSTEM Connections - **Location:** Ontario, CA, United States - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Information Engineering, Distributed Computing Environment, Distributed Systems, Machine Learning, Build Management, Kubernetes, Machine Learning Operations, Data Pipelines - **Published:** June 24, 2026 - **Apply:** https://www.dice.com/job-detail/37bcc396-4fb8-4ea7-87ae-8f410fe2c2a6 ## About the Role Experience Baseline: Minimum of five (5) to ten (10) years of experience in designing, building, and maintaining large-scale ML infrastructure or distributed systems. Infrastructure Mastery: Deep expertise in container orchestration (e.g., Kubernetes), distributed training, and cloud-native infrastructure. Pipeline Expertise: Proven track record of managing massive-scale data pipelines and feature stores. Collaboration: Strong "bridge-builder" personality-you are comfortable working in a high-velocity environment alongside both pure research scientists and core software engineers. Preferred Attributes Proven background at large-scale AI/ML-driven technology companies. Experience with foundational model training infrastructure and techniques (e.g., model distillation, fine-tuning at scale). A "Scalability Mindset"-you prioritize modular, reusable, and testable code over "quick and dirty" scripts. ## Description ML Infrastructure & Automation Workflow Architecture: Design and build end-to-end ML workflows and automated pipelines that minimize manual intervention and accelerate the path from experimentation to production. Training & Serving Platforms: Architect and scale our distributed training and model-serving infrastructure. Build the platforms that handle foundational model training, knowledge distillation, and high-performance inference. Data Engineering at Scale: Develop robust data sampling and feature generation platforms that provide high-quality input for our ML systems. Automation & Reliability: Build foundational tools that standardize how we train, track, and deploy models, ensuring high platform reliability and minimal deployment drift. Performance & Optimization Cost & Efficiency: Drive architectural decisions that optimize our infrastructure footprint. Implement smart resource management and cost-optimization strategies for large-scale training clusters. Developer Productivity: Build "developer-first" internal tools that reduce the cognitive load on researchers, allowing them to focus on model logic rather than infrastructure configuration. ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [The state of MLOps - machine learning in production at enterprise scale](https://www.wearedevelopers.com/videos/369-the-state-of-mlops-machine-learning-in-production-at-enterprise-scale) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)