Ai/Ml Platform Cloud Infrastructure Engineer - Lial Product.
Role details
Job location
Tech stack
Job description
With over 20 years of experience, our global network of passionate technologists and pioneering craftsmen deliver cutting-edge technology and game-changing consulting to companies on the brink of transformation.Since ****, we have grown from a Java company into a full-service digital consulting company with ****+ professionals working on a worldwide ambition.We are organized in complementary chapters - teams with a tremendous amount of knowledge and experience within a particular field, such as Agile, DevOps, Data and AI, Cloud, Software Technology, Functional Programming, Low Code, and Microsoft.We help the world's top 250 companies and category leaders overcome digital challenges, embrace innovation, adopt new technology, and implement new business models.In addition to high-quality consulting, we also provide offshoring and nearshoring services.For more details please visit Xebia, we put 'People First'-committed to attracting diverse talent and fostering an inclusive, respectful workplace where everyone is valued for their contributions.We welcome all individuals and evaluate solely on the quality of their work and teamwork.Role OverviewWe are seeking experienced AI/ML Platform ML Engineers to support the development and operationalization of the LIAL product.This role combines strong machine learning expertise with production-grade software engineering to build, scale, and maintain robust ML platform capabilities.You will be responsible for designing and implementing end-to-end ML workflows-from experimentation and training through to deployment, monitoring, and retraining-ensuring all models are production-ready, reproducible, and governed by best practices.Key ResponsibilitiesApply strong software engineering discipline to ML development, transforming exploratory notebooks into modular, reusable, and testable Python packages .Design and enforce clear interface contracts for ML components to support maintainability and scalability.Implement experiment tracking frameworks (e.G., MLflow or Vertex AI Experiments), ensuring:Full capture of parameters, metrics, artifacts, and dataset lineageReproducibility of results from a commit hash alonePromote best practices for code versioning, testing, and documentation across ML workflows.Training, Evaluation & Hyperparameter OptimizationDesign and implement distributed training pipelines (multi-GPU / multi-node), ensuring:Robust checkpointingFault tolerance and recoverabilityDevelop standardized evaluation templates that include:Core performance metricsBias and fairness assessmentsShadow-mode testing against baseline modelsMove beyond simple validation by ensuring models are evaluated under realistic production scenarios .Model Packaging, Serving & DeploymentBuild and maintain standardized model packaging templates, including:TensorFlow SavedModelTorchScriptONNXCreate versioned, production-ready serving containers and publish them to registries (e.G., Artifact Registry).Canary releases with traffic splittingSafe rollback proceduresBoth real-time (online) and batch inference use casesLeverage platforms such as Vertex AI Endpoints to operationalize model serving.Production Monitoring & RetrainingDesign templates covering the entire post-deployment lifecycle, including:Prediction quality monitoringClearly defined alert thresholdsBuild and maintain automated retraining pipelines triggered by monitoring signals.Define and enforce model lifecycle governance, including:Operational runbooksEnsure no model operates in production without observability and traceability .Ways of Working & EngagementWork independently and autonomously, owning deliverables end-to-end.Apply a security-first mindset in all platform and ML engineering activities.Demonstrate strong:Planning and prioritization skillsCommunication and stakeholder engagementReporting and documentation disciplineCollaborate effectively within cross-functional teams including product, data, and platform engineering.Required Experience & SkillsProven experience as an ML Engineer in production environments (not purely research-focused).Strong proficiency in Python and modern ML frameworks (TensorFlow, PyTorch).Hands-on experience with:ML lifecycle tooling (MLflow, Vertex AI, or equivalent)Distributed training and scalable compute environmentsContainerization (Docker) and deployment pipelinesExperience with cloud-native ML platforms, preferably Google Cloud / Vertex AI .Solid understanding of:Model evaluation beyond accuracy (fairness, robustness, monitoring)CI/CD for ML systems (MLOps practices)Familiarity with artifact management and version control systems .Nice to HaveExperience building enterprise AI/ML platforms supporting multiple teams/productsKnowledge of data governance, lineage, and compliance frameworksExposure to high-scale ML systems and real-time inference architecturesExperience implementing automated retraining and adaptive learning systemsEngagement ModelFocus: Delivery of AI/ML Platform Engineering capabilities to support LIAL product developmentWorking Style: Autonomous delivery with structured reporting and stakeholder alignmentSuccess CriteriaReproducible ML pipelines with full traceabilityProduction-grade deployment patterns with zero-downtime releasesRobust monitoring and automated retraining pipelines in placeStandardized templates enabling scalable ML development across teamsCompensationSalary Range: €68,000 - €82,000 gross per year, depending on experience, skills, and overall fit for the role.#J-*****-Ljbffr
Requirements
Ways of Working & EngagementWork independently and autonomously, owning deliverables end-to-end.Apply a security-first mindset in all platform and ML engineering activities.Demonstrate strong:Planning and prioritization skillsCommunication and stakeholder engagementReporting and documentation disciplineCollaborate effectively within cross-functional teams including product, data, and platform engineering.Required Experience & SkillsProven experience as an ML Engineer in production environments (not purely research-focused). Strong proficiency in Python and modern ML frameworks (TensorFlow, PyTorch). Hands-on experience with:ML lifecycle tooling (MLflow, Vertex AI, or equivalent)Distributed training and scalable compute environmentsContainerization (Docker) and deployment pipelinesExperience with cloud-native ML platforms, preferably Google Cloud / Vertex AI . Solid understanding of:Model evaluation beyond accuracy (fairness, robustness, monitoring)CI/CD for ML systems (MLOps practices)Familiarity with artifact management and version control systems . Nice to HaveExperience building enterprise AI/ML platforms supporting multiple teams/productsKnowledge of data governance, lineage, and compliance frameworksExposure to high-scale ML systems and real-time inference architecturesExperience implementing automated retraining and adaptive learning systemsEngagement ModelFocus: Delivery of AI/ML Platform Engineering capabilities to support LIAL product developmentWorking Style: Autonomous delivery with structured reporting and stakeholder alignmentSuccess CriteriaReproducible ML pipelines with full traceabilityProduction-grade deployment patterns with zero-downtime releasesRobust monitoring and automated retraining pipelines in placeStandardized templates enabling scalable ML development across teamsCompensationSalary Range: €68,000 - €82,000 gross per year, depending on experience, skills, and overall fit for the role. #J-*****-Ljbffr