> Markdown version of [/jobs/ext/2589228-foundational-staff-machine-learning-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2589228-foundational-staff-machine-learning-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # foundational Staff Machine Learning Infrastructure Engineer - **Company:** ATOM, INC. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $224,000.0 - $280,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Artificial Neural Networks, Cloud Computing, Computer Clusters, Computer Programming, Continuous Integration, Information Engineering, Distributed Data Store, Distributed Systems, Python (Programming Language), Machine Learning, Metadata, Network Architecture, Software Engineering, System Programming, Systems Integration, Scripting, Graphics Processing Unit (GPU), Model Validation, Backend, Integration Tests, Kubernetes, Data Management, Machine Learning Operations, Multiplatform, Data Pipelines, Golang - **Published:** August 31, 2026 - **Apply:** https://www.careerbuilder.com/job-details/staff-machine-learning-infrastructure-engineer-san-francisco-ca--66274b75-2d54-4f04-bcd1-494e7669c62a ## About the Role * 8+ years of professional software engineering career experience * Strong backend systems programming skills with proficiency in Go, Python, Java or similar (with familiarity or exposure to Rust considered a plus). * Proficiency with Kubernetes for container orchestration and building cloud-agnostic environments from scratch. * Experience implementing distributed ML compute frameworks (e.g., Ray) to coordinate large pools of GPUs for heavy, multi-node workloads. * Hands-on experience building MLOps pipelines, metadata tracking architectures, and model registries using platforms like MLflow. * Prior experience managing high-throughput data pipelines using modern distributed data engines to feed data-hungry neural network architectures., Artificial Intelligence (AI), Automation, Cloud Computing, Compensation Management, Computer Programming, Continuous Integration, Cross-Functional, Data Management, Disability Insurance, Distributed Computing, GPU (Graphics Processing Unit), Hardware-Software Integration, High Throughput, Integration Testing, Java, Life Insurance, Machine Learning, Manufacturing, Metadata, Mineral/Metal Mining, Model Validation, Multiplatform/Cross-Platform, Network Architecture/Engineering, Neural Networks, Performance Metrics, Product Development, Python Programming/Scripting Language, Real Estate, Robotics, Software Engineering, Systems/Internals Programming, Transportation Modeling, Vehicle Fleets ## Description We are seeking a foundational Staff Machine Learning Infrastructure Engineer to design and build the large-scale ML training infrastructure that powers our next-generation autonomous transport models. In this role, you will design the high-performance training pipelines and validation environments that enable our world-class robotics and ML researchers to iterate rapidly. You will own the challenge of scaling distributed GPU workloads to support a high volume of concurrent training runs across an expanding vehicle fleet, building a platform that can flexibly run on whatever GPU capacity is available, regardless of provider or environment, directly accelerating innovation across the platform. * Training Infrastructure: Design, implement, and scale repeatable machine learning infrastructure utilizing Kubernetes to support large-scale distributed GPU training of novel neural networks. * Distributed Computing & Orchestration: Leverage distributed compute frameworks to efficiently manage and execute a high volume of complex ML training jobs concurrently across large GPU clusters. * Experiment Tracking & MLOps: Integrate advanced model management and experiment tracking tools to provide researchers with deep observability into training metrics and run performance. * Data Engineering Pipelines: Build and optimize high-throughput data ingestion pipelines to seamlessly stream petabyte-scale multi-sensor vehicle logs into training environments. * Validation at Scale: Architect robust infrastructure for autonomous model validation and continuous integration testing, ensuring new vehicle policy releases are entirely regression-free. * Cross-Functional Collaboration: Partner closely with core robotics engineers and machine learning researchers to eliminate workflow bottlenecks and accelerate the deploy-to-vehicle lifecycle. ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [A Data Mesh needs Open Metadata](https://www.wearedevelopers.com/videos/505-a-data-mesh-needs-open-metadata) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [The state of MLOps - machine learning in production at enterprise scale](https://www.wearedevelopers.com/videos/369-the-state-of-mlops-machine-learning-in-production-at-enterprise-scale) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)