> Markdown version of [/jobs/ext/2825812-ml-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2825812-ml-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Infrastructure Engineer - **Company:** Zipline - **Location:** South San Francisco, CA, United States - **Experience:** Expert - **Salary:** $160,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Cloud Computing, Data Files, Data Infrastructure, Programming Tools, Distributed Computing Environment, Python (Programming Language), Machine Learning, Cloud Services, Tensorflow, Software Engineering, Software Systems, Data Ingestion, Pytorch, Deep Learning, Cloudformation, Kubernetes, Infrastructure Automation Frameworks, Performance Monitor, Data Management, Machine Learning Operations, Terraform, Data Pipelines - **Published:** September 10, 2026 - **Apply:** https://startup.jobs/ml-infrastructure-engineer-zipline-9981794 ## About the Role * 3+ years of professional software engineering experience, ideally including ML infrastructure, data infrastructure, robotics, autonomy, aerospace, medical devices, or another safety-critical hardware/product environment. * Strong software engineering practices in Python in a production setting; comfort designing APIs, services, schemas, jobs, and operational workflows. * Experience building reproducible data pipelines and machine-learning pipelines. * Experience monitoring data statistics, system performance metrics, pipeline failures, and model/evaluation signals. * Working knowledge of ML concepts such as datasets, training, evaluation, optimization, statistics, and modern deep learning workflows. * Generalist mindset and willingness to work across cloud services, data platforms, developer tooling, and embedded/robotics-adjacent constraints. * Experience with PyTorch or similar ML frameworks. * Strong ownership, clear communication, and interest in building secure systems for mission-critical workflows. * Experience with Kubernetes or other container orchestration systems for production workloads. * Experience with cloud and on-premise production infrastructure, preferably AWS, and infrastructure-as-code tools such as Terraform or CloudFormation., * Experience deploying or evaluating ML systems on real robots, autonomous vehicles, drones, or other hardware products. * Experience with large-scale training systems, feature stores, data/versioned artifact stores, model registries, or experiment tracking. * Experience with annotation systems, dataset inspection tooling, or active-learning workflows. ## Description As an ML Training & Inference Infrastructure Engineer on the Data Platform team you will be building and scaling the systems powering our data flywheel. This person will work at the intersection of autonomy and the infrastructure, owning systems that make ML development faster, reproducible, observable, and safe. This role is for a strong software engineer who enjoys the full ML development cycle: data ingestion, processing pipelines, dataset management, distributed training, continuous model integration, evaluation, and deployment. The ideal candidate has strong production engineering habits and is excited to build infrastructure that helps real autonomous systems improve over time. What You'll Do * Build and operate software infrastructure that enables learning algorithms to leverage Zipline's large-scale (quickly growing!) fleet data. * Design scalable, maintainable data and ML infrastructure for autonomy teams, including dataset creation, validation, training, evaluation, and deployment. * Own and improve data pipelines that feed into the ML development loop. * Identify and mitigate bottlenecks in the ML development cycle, especially around orchestration, performance, and reproducibility to increase the rate at which we can improve and scale the delivery experience. ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Introduction to TXT](https://www.wearedevelopers.com/videos/30-introduction-to-txt) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Industrializing your Data Science capabilities](https://www.wearedevelopers.com/videos/178-industrializing-your-data-science-capabilities) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)