Ai/Ml Platform Cloud Infrastructure Engineer - Lial Product.

Xebia
Málaga, Spain
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
€68,000.0 - €82,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Agile Methodology Artificial Intelligence Cloud Computing Cloud Engineering Continuous Integration Data Governance DevOps Distributed Computing Environment Fault Tolerance Python (Programming Language) Machine Learning
+15 more
Tensorflow Azure Machine Learning Software Engineering Management of Software Versions Google Cloud Pytorch Delivery Pipeline Model Validation Containerization Templating ONNX (Open Neural Network Exchange) Format Machine Learning Operations Functional Programming Software Version Control Docker

Job description

With over 20 years of experience, our global network of passionate technologists and pioneering craftsmen deliver cutting-edge technology and game-changing consulting to companies on the brink of transformation. Since **, we have grown from a Java company into a full-service digital consulting company with **+ professionals working on a worldwide ambition.We are organized in complementary chapters - teams with a tremendous amount of knowledge and experience within a particular field, such as Agile, DevOps, Data and AI, Cloud, Software Technology, Functional Programming, Low Code, and Microsoft.We help the world’s top 250 companies and category leaders overcome digital challenges, embrace innovation, adopt new technology, and implement new business models. In addition to high-quality consulting, we also provide offshoring and nearshoring services.For more details please visitAt Xebia, we put ‘People First’-committed to attracting diverse talent and fostering an inclusive, respectful workplace where everyone is valued for their contributions. We welcome all individuals and evaluate solely on the quality of their work and teamwork.Role OverviewWe are seekingexperienced AI/ML Platform ML Engineersto support the development and operationalization of theLIAL product. This role combines strong machine learning expertise with production-grade software engineering to build, scale, and maintain robust ML platform capabilities.You will be responsible for designing and implementing end-to-end ML workflows-from experimentation and training through to deployment, monitoring, and retraining-ensuring all models are production-ready, reproducible, and governed by best practices.Key ResponsibilitiesApply strongsoftware engineering discipline to ML development, transforming exploratory notebooks intomodular, reusable, and testable Python packages.Design and enforceclear interface contractsfor ML components to support maintainability and scalability.Implementexperiment tracking frameworks(e.g., MLflow or Vertex AI Experiments), ensuring:Full capture of parameters, metrics, artifacts, and dataset lineageReproducibility of results from acommit hash alonePromote best practices forcode versioning, testing, and documentationacross ML workflows.Training, Evaluation & Hyperparameter OptimizationDesign and implementdistributed training pipelines(multi-GPU / multi-node), ensuring:Robust checkpointingFault tolerance and recoverabilityDevelop standardizedevaluation templatesthat include:Core performance metricsBias and fairness assessmentsShadow-mode testing against baseline modelsMove beyond simple validation by ensuring models are evaluated underrealistic production scenarios.Model Packaging, Serving & DeploymentBuild and maintainstandardized model packaging templates, including:TensorFlow SavedModelTorchScriptONNXCreateversioned, production-ready serving containersand publish them to registries (e.g., Artifact Registry).Canary releases with traffic splittingSafe rollback proceduresBothreal-time (online)andbatch inferenceuse casesLeverage platforms such asVertex AI Endpointsto operationalize model serving.Production Monitoring & RetrainingDesign templates covering theentire post-deployment lifecycle, including:Prediction quality monitoringClearly defined alert thresholdsBuild and maintainautomated retraining pipelinestriggered by monitoring signals.Define and enforcemodel lifecycle governance, including:Operational runbooksEnsure no model operates in production withoutobservability and traceability.Ways of Working & EngagementWorkindependently and autonomously, owning deliverables end-to-end.Apply asecurity-first mindsetin all platform and ML engineering activities.Demonstrate strong:Planning and prioritization skillsCommunication and stakeholder engagementReporting and documentation disciplineCollaborate effectively within cross-functional teams including product, data, and platform engineering.Required Experience & SkillsProven experience as anML Engineer in production environments(not purely research-focused).Strong proficiency inPythonand modern ML frameworks (TensorFlow, PyTorch).Hands-on experience with:ML lifecycle tooling(MLflow, Vertex AI, or equivalent)Distributed training and scalable compute environmentsContainerization (Docker) and deployment pipelinesExperience withcloud-native ML platforms, preferablyGoogle Cloud / Vertex AI.Solid understanding of:Model evaluation beyond accuracy (fairness, robustness, monitoring)CI/CD for ML systems (MLOps practices)Familiarity withartifact management and version control systems.Nice to HaveExperience buildingenterprise AI/ML platformssupporting multiple teams/productsKnowledge ofdata governance, lineage, and compliance frameworksExposure tohigh-scale ML systemsand real-time inference architecturesExperience implementingautomated retraining and adaptive learning systemsEngagement ModelFocus: Delivery ofAI/ML Platform Engineering capabilitiesto supportLIAL product developmentWorking Style: Autonomous delivery with structured reporting and stakeholder alignmentSuccess CriteriaReproducible ML pipelines with full traceabilityProduction-grade deployment patterns with zero-downtime releasesRobust monitoring and automated retraining pipelines in placeStandardized templates enabling scalable ML development across teamsCompensationSalary Range:€68,000 - €82,000 gross per year, depending on experience, skills, and overall fit for the role.#J-*****-Ljbffr

Requirements

Planning and prioritization skills Communication and stakeholder engagement Reporting and documentation discipline Collaborate effectively within cross-functional teams including product, data, and platform engineering. Required Experience & Skills Proven experience as an ML Engineer in production environments (not purely research-focused). Strong proficiency in Python and modern ML frameworks (TensorFlow, PyTorch). Hands-on experience with: ML lifecycle tooling (MLflow, Vertex AI, or equivalent) Distributed training and scalable compute environments Containerization (Docker) and deployment pipelines Experience with cloud-native ML platforms , preferably Google Cloud / Vertex AI . Solid understanding of: Model evaluation beyond accuracy (fairness, robustness, monitoring) CI/CD for ML systems (MLOps practices) Familiarity with artifact management and version control systems

Benefits & conditions

Salary Range: €68,000 - €82,000 gross per year, depending on experience, skills, and overall fit for the role. #J-*****-Ljbffr

About the company

Málaga, España

With over 20 years of experience, our global network of passionate technologists and pioneering craftsmen deliver cutting-edge technology and game-changing consulting to companies on the brink of transformation. Since **, we have grown from a Java company into a full-service digital consulting company with **+ professionals working on a worldwide ambition. We are organized in complementary chapters - teams with a tremendous amount of knowledge and experience within a particular field, such as Agile, DevOps, Data and AI, Cloud, Software Technology, Functional Programming, Low Code, and Microsoft. We help the world’s top 250 companies and category leaders overcome digital challenges, embrace innovation, adopt new technology, and implement new business models. In addition to high-quality consulting, we also provide offshoring and nearshoring services.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

Videos

See all

Related articles

See all