> Markdown version of [/jobs/ext/1231641-ai-ml-platform-cloud-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1231641-ai-ml-platform-cloud-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI/ML Platform Cloud Infrastructure Engineer - **Company:** Xebia - **Location:** Jávea, Spain - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Engineering, Distributed Computing Environment, Fault Tolerance, Python (Programming Language), Machine Learning, Tensorflow, Azure Machine Learning, Software Engineering, Google Cloud, Pytorch, Delivery Pipeline, Containerization, Templating, ONNX (Open Neural Network Exchange) Format, Machine Learning Operations, Docker - **Published:** July 11, 2026 - **Apply:** https://es.trabajo.org/oferta-5354-c6d36ae79c1423d442455d96efd1b53a ## About the Role cross-functional teams including product, data, and platform engineering. Required Experience & Skills Proven experience as an ML Engineer in production environments (not purely research-focused). Strong proficiency in Python and modern ML frameworks (TensorFlow, PyTorch). Hands-on experience with: ML lifecycle tooling (MLflow, Vertex AI, or equivalent) Distributed training and scalable compute environments Containerization (Docker) and deployment pipelines Experience with cloud-native ML platforms , preferably Google Cloud / Vertex AI . S ## Description respectful workplace where everyone is valued for their contributions. We welcome all individuals and evaluate solely on the quality of their work and teamwork. Role Overview We are seeking experienced AI/ML Platform ML Engineers to support the development and operationalization of the LIAL product . This role combines strong machine learning expertise with production-grade software engineering to build, scale, and maintain robust ML platform capabilities. You will be responsible for designing and implementing end-to-end ML workflows-from experimentation and training through to deployment, monitoring, and retraining-ensuring all models are production-ready, reproducible, and governed by best practices. Key Responsibilities Apply strong software engineering discipline to ML development , transforming exploratory notebooks into modular, reusable, and testable Python packages . Design and enforce clear interface contracts for ML components to support maintainability and scalability. Implement experiment tracking frameworks (e.g., MLflow or Vertex AI Experiments), ensuring: Full capture of parameters, metrics, artifacts, and dataset lineage Reproducibility of results from a commit hash alone Promote best practices for code versioning, testing, and documentation across ML workflows. Training, Evaluation & Hyperparameter Optimization Design and implement distributed training pipelines (multi-GPU / multi-node), ensuring: Robust checkpointing Fault tolerance and recoverability Develop standardized evaluation templates that include: Core performance metrics Bias and fairness assessments Shadow-mode testing against baseline models Move beyond simple validation by ensuring models are evaluated under realistic production scenarios . Model Packaging, Serving & Deployment Build and maintain standardized model packaging templates , including: TensorFlow SavedModel TorchScript ONNX Create versioned, production-ready serving containers and publish them to registries (e.g., Artifact Registry). Canary releases with traffic splitting Safe rollback procedures Both real-time (online) and batch inference use cases Leverage platforms such as Vertex AI Endpoints to operationalize model serving. Production Monitoring & Retraining Design templates covering the entire post-deployment lifecycle , including: Prediction quality monitoring Clearly defined alert thresholds Build and maintain automated retraining pipelines triggered by monitoring signals. Define and enforce model lifecycle governance , including: Operational runbooks Ensure no model operates in production without observability and traceability . Ways of Working & Engagement Work independently and autonomously , owning deliverables end-to-end. Apply a security-first mindset in all platform and ML engineering activities. Demonstrate strong: Planning and prioritization skills Communication and stakeholder engagement Reporting and documentation discipline Collaborate effectively within ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Backstage Software Templates for Java Developers](https://www.wearedevelopers.com/videos/1596-backstage-software-templates-for-java-developers) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)