> Markdown version of [/jobs/ext/2857656-ml-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2857656-ml-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Platform Engineer - **Company:** Agile Robots Ag - **Location:** München, Germany - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Cloud Computing, Continuous Integration, Distributed Computing Environment, Monitoring of Systems, Python (Programming Language), Prometheus, Azure Machine Learning, Software Engineering, Software Systems, Workflow Management Systems, Pytorch, Grafana, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Slurm, Machine Learning Operations, Terraform, Software Version Control, Docker - **Published:** September 12, 2026 - **Apply:** https://www.adzuna.de/details/5879926325 ## About the Role * Background and Experience: Degree in Computer Science, Software Engineering, or a related field, with professional experience building and operating ML or software infrastructure in production. * Distributed Training: Experience designing and operating distributed training systems on Kubernetes and Docker, using PyTorch Distributed, DeepSpeed, and schedulers such as SLURM. * CI/CD for ML: Experience building CI/CD pipelines that support reliable model testing, training, and deployment. * Cloud Infrastructure: Experience operating ML workloads on cloud infrastructure, preferably AWS. * Experiment Tracking: Hands-on experience with experiment tracking and model versioning using tools such as MLflow or Weights & Biases. * Observability: Experience with monitoring and drift detection using tools such as Prometheus and Grafana. * Software Engineering: Python and system design skills, with experience building and operating ML systems beyond the prototype stage. Beneficial Skills * Multimodal Systems: Experience with large-scale or multimodal ML systems such as vision-language-action models. * Infrastructure As Code: Familiarity with infrastructure-as-code tools such as Terraform. * ML Orchestration: Experience with ML pipeline and orchestration tools. * Distributed Compute: Exposure to high-performance or distributed compute environments. ## Description * Training Infrastructure: Design and scale distributed training workflows for large models using tools such as PyTorch Distributed, DeepSpeed, and cluster schedulers like SLURM or Kubernetes. * ML Platform: Build and maintain containerised ML environments that support reproducible experimentation and benchmarking. * CI/CD Pipelines: Develop and maintain CI/CD pipelines for machine learning systems to enable reliable testing, training, and deployment of models. * Lifecycle Management: Implement experiment tracking, model versioning, and reproducibility workflows using tools such as ClearML or Weights & Biases. * Observability: Set up monitoring systems such as Prometheus and Grafana to track model performance and system health and detect drift in production. * Cross-Team Collaboration: Work with research, data, and robotics teams to connect new models to robust production systems., * A dynamic high-tech company, combined with financial soundness and world-class investors. * Join an interdisciplinary, international team with 60+ different nationalities in a collaborative work environment. * Lots of development opportunities as we continue to grow. * Challenging tasks and impactful projects alongside experts that enable professional and personal growth. * Corporate Benefits Program covering health, mobility, and learning for 100€ net per month. * Modern office facilities with a rooftop terrace overlooking Munich, free drinks & fruits, and regular company events contribute to a good working environment ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Best Coding Boot Camps in Germany](https://www.wearedevelopers.com/magazine/237-best-coding-boot-camps-in-germany) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)