> Markdown version of [/jobs/ext/1306663-staff-lead-python-engineer-ml-platform-ops](https://www.wearedevelopers.com/jobs/ext/1306663-staff-lead-python-engineer-ml-platform-ops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff/Lead Python Engineer (ML Platform/Ops) - **Company:** Tipo De - **Location:** Bellprat, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Airflow, C++ (Programming Language), Computer Programming, Continuous Integration, Linux, Distributed Systems, Python (Programming Language), Machine Learning, Site Reliability Engineering Practices, Mesos, Azure Machine Learning, Data Streaming, Workflow Management Systems, Concurrency, Build Management, Kubernetes, Apache Kafka, Machine Learning Operations, Hardware Infrastructure - **Published:** July 17, 2026 - **Apply:** https://www.adzuna.es/contact-us.html ## About the Role Requirements:5+ years building distributed systems; 3+ years in MLOps or ML platform engineeringStrong knowledge of Linux/OS internals, networking, concurrency, and performance profilingDeep expertise in Kubernetes (bonus: Mesos) and GPU infrastructure managementProficiency in high-performance programming (Java, Rust, Go, C++; strong Python skills)Experience designing and operating production model platforms (registry, training, serving, monitoring)Proven experience leading technical teams and implementing organization-wide platform solutionsFamiliarity with CI/CD, SRE practices, observability, and reliability enablementStrong collaboration, mentoring, and communication skillsPreferred:Experience with streaming/workflow tools (Kafka, Argo, Temporal, Airflow)Hands-on work with eBPF observability, perf tooling, or io_uringExpertise in ML/AI cost optimization, multi-tenant quotas, and fairnessExperience authoring Golden Paths (service templates, CI/CD blueprints ## Description Inscribirse en esta oferta This position is posted by Jobgether on behalf of a partner company. We are currently looking for a Core & ML Ops Team Lead in Spain.This role is ideal for an experienced technical leader in MLOps and distributed systems, responsible for building and maintaining the scalable infrastructure that supports mission-critical services. You will lead a cross-functional team in designing platforms for model training, orchestration, deployment, and monitoring while ensuring high performance, reliability, and security. The position combines hands-on engineering with strategic team leadership, driving adoption of best practices, automation, and observability across the organization. You will collaborate with product, operations, and security teams to implement robust platforms that empower engineers to build and deploy services confidently. Mentorship, knowledge sharing, and establishing production-ready standards are central to your impact. This role allows you to shape platform strategy while staying deeply engaged in cutting-edge technologies and ML operations at scale.Accountabilities: + Lead the Core & MLOps team, overseeing roadmap, prioritization, delivery, and mentoring + Design, develop, and maintain scalable infrastructure for model training, serving, and monitoring + Build and maintain the Golden Path: reference repositories, scaffold CLIs, CI/CD pipelines, runtime contracts, and production-ready defaults + Operate secure, multi-tenant model registries and orchestration platforms with standardized experiment and evaluation frameworks + Integrate AI/ML capabilities as managed platform services with cost and governance controls + Collaborate with product engineering, operations, and security teams on adoption and rollout plans + Promote best practices in observability, reliability, cost governance, and platform standardization ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)