> Markdown version of [/jobs/ext/3584532-central-software-csw-ml-platform-team](https://www.wearedevelopers.com/jobs/ext/3584532-central-software-csw-ml-platform-team). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Central Software (CSW) ML Platform team - **Company:** Boston Dynamics - **Location:** United States - **Experience:** Expert - **Salary:** $150,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Airflow, Amazon Web Services, Computer Clusters, Continuous Integration, Extract Transform Load (ETL), Python (Programming Language), Networking Basics, Node.Js, NumPy, Scrum Methodology, Ansible, Tensorflow, Azure Machine Learning, TypeScript, Data Processing, Pytorch, Evaluation Pipelines, Database Optimization, Apache Spark, Indexer, Amazon Virtual Private Cloud (VPC), Pandas, Kubernetes, Information Technology, Low Latency, Web Technologies, Slurm, Machine Learning Operations, Terraform, Docker - **Published:** October 4, 2026 - **Apply:** https://startup.jobs/staff-ml-ops-engineer-boston-dynamics-inc-10275365 ## About the Role * 5+ years of experience as a Senior Software Engineer or ML Engineer * Proficiency in Python and related ML frameworks (PyTorch, TensorFlow, Pandas, NumPy) * Experience with cloud platforms (e.g., GCP, AWS) and scalable ML deployment methods (Docker, Kubernetes, Ansible, Terraform) * Experience with GPU cluster management and scheduling (e.g., Slurm, Kueue, or similar) * Experience with CI/CD practices applied to ML pipelines * Experience with experiment tracking and model/data versioning tools (e.g., MLflow, Weights & Biases, DVC) * Experience building scalable data and ETL pipelines (e.g., Spark, Airflow) alongside data processing, augmentation, and cleaning techniques. * Experience with Agile, Scrum, or other lean methodologies; ability to work collaboratively in cross-functional teams * Bachelor's degree in Engineering, Computer Science, or a related technical field, or equivalent practical experience Nice to have * Experience with networking fundamentals such as IAP, Tailscale, Shared VPC, NAT * Familiarity with database optimization concepts such as indexing and connection pooling. * Experience with TypeScript, Node, and related full-stack web technologies to build internal tooling, visualization dashboards, and web-based interfaces for MLOps platforms * Experience with on-robot/edge deployment constraints (latency, compute limits, OTA model updates * Experience with annotation tools such as SAM or Co-Tracker We are interested in all qualified candidates eligible to work in the United States. However, we are not able to sponsor visas for this position. ## Description * Transform proofs of concept into scalable solutions, helping product teams deliver new robot capabilities to customers * Evolve and scale fielded solutions, enabling continuous model improvement and redeployment * Work with stakeholders across BD to understand requirements, ensuring deployed solutions meet end-user needs * Own end-to-end delivery of new capabilities, spanning implementation, testing, deployment, and operations * Maintain our GPU clusters and develop automation to monitor and improve cluster health * Use observability tools to monitor system health and root-cause problems across the platform * Drive accuracy and efficiency by profiling and optimizing ML data, training, and evaluation pipelines. * Work closely with other members of the ML Platform team to implement, deploy, and maintain ML infrastructure * Be an active participant in our agile development process, coordinating work with others, calling out challenges, and regularly communicating progress * Use your experience to mentor and upskill peers and other contributors across the organization ## Related Videos - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [How to implement convenient Python bindings to C++](https://www.wearedevelopers.com/videos/618-how-to-implement-convenient-python-bindings-to-c) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)