Mlops Platform Architect
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
Role As an MLOps / ML Platform Architect , you own the path from model prototype to reliable production on Azure and Databricks .You design and operate the systems that support training, deployment, monitoring, governance, and retraining, ensuring that AI and ML solutions remain reliable, scalable, secure, and trustworthy in production.You combine hands-on MLOps engineering with solution architecture.You can assess workloads, identify bottlenecks, design the target approach, and turn technical requirements into an actionable delivery proposal.You work closely with the AI Solutions Architect, AI Engineers, Data Engineers, and platform teams, and you are accountable for the quality of AI/ML delivery.The client cloud landing zone and core data platform are typically already in place.Your role is to integrate with and build on that foundation rather than design it from scratch.Data Engineering owns ingestion and curation through the silver or gold layer; MLOps ownership begins with feature pipelines and the production ML lifecycle built on top of those curated datasets.Your responsibilities will include: Design and operate ML platforms on Azure Databricks, including feature pipelines and feature management, MLflow experiment tracking, model registry, training pipelines, and batch or online serving.Define and implement model lifecycle standards covering versioning, reproducibility, evaluation gates, promotion, deployment, rollback, and retraining.Build production-ready deployment patterns, including models exposed through APIs, automated MLflow workflows, and CI/CD pipelines for models and data.Own production monitoring for ML systems, including data and concept drift, model quality, latency, reliability, and cost, with appropriate alerting and retraining triggers.Design scalable Databricks workloads, identify architectural and performance bottlenecks, and translate findings into clear solution designs and delivery proposals.Lead decisions on ML tooling, integration patterns, and platform evolution, with Databricks as the primary platform and Azure ML used where appropriate.Build reusable templates, reference implementations, and accelerators that enable AI Engineers to move solutions into production independently.Establish ML governance practices, including lineage, audit trails, model documentation, access controls, and compliance-ready operating processes.Support GenAI delivery when needed, including deployment and monitoring of RAG and LLM applications using Azure OpenAI or Azure AI Foundry, vector search, evaluation, and guardrails.Apply the client’s existing cloud, security, networking, and operating standards to ML systems and integrate with the established landing zone and data platform.Ensure ML platforms and workloads are reliable, cost-efficient, secure, maintainable, and observable.Partner with engineering teams and senior stakeholders to guide implementation, mentor contributors, and accelerate adoption.Requirements To succeed in this role, you bring a combination of expertise, experience, and skills including: 5+ years of experience in MLOps, ML platform engineering, or a similar role , with significant hands-on experience running ML systems in production.Proven experience taking models to production and keeping them reliable , with concrete examples covering deployment, monitoring, incident resolution, and lifecycle management.Deep hands-on expertise in Azure and Databricks , especially Azure Databricks, MLflow, and Unity Catalog.Experience with Azure ML is a strong advantage.Strong experience designing Databricks solutions, diagnosing workload bottlenecks, and producing clear technical designs and implementation proposals.Practical Python software-engineering skills for MLOps and platform work, including packaging, testing, APIs, and maintainable automation.Strong knowledge of containers and model-serving patterns.Experience with Kubernetes or AKS is beneficial but not required.Deep experience building CI/CD pipelines for models and data, beyond infrastructure and application deployment alone.Confidence with Terraform and Infrastructure as Code, particularly when integrating ML services into an existing Azure landing zone.Working knowledge of cloud networking and security, including IAM, encryption, secrets management, governance, and observability.Understanding of ML governance, lineage, auditability, and model documentation.A scientific or life-sciences background is not required.Working knowledge of GenAI technologies and patterns, including Azure OpenAI or Azure AI Foundry, RAG, vector search, evaluation, and guardrails.Ability to influence senior stakeholders, communicate technical strategy clearly, and collaborate effectively across AI, data, platform, security, and product teams.Strong mentoring skills and a pragmatic approach to helping engineering teams deliver independently.From the outset, you can independently: Deploy a model behind a production-ready API or managed serving endpoint.Configure MLflow experiment tracking and model registration.Build or improve a CI/CD pipeline for models and data.Assess an ML workload, identify bottlenecks, and produce a sound solution design and delivery proposal.Over time, you establish repeatable standards and platform capabilities that make ML delivery faster, safer, and more reliable across teams.Benefits What we offer A competitive compensation package A yearly education budget to steep your learning curve A yearly sport budget because a fit body leads to a fit mind A flexible working culture because your work-life balance matters to us A position that enables you to have an impact on 1’000s of people, and the whole company’s growth.An international, knowledgeable, and passionate team with a strong collaborative mindset
Requirements
Requirements To succeed in this role, you bring a combination of expertise, experience, and skills including: 5+ years of experience in MLOps, ML platform engineering, or a similar role , with significant hands-on experience running ML systems in production. Proven experience taking models to production and keeping them reliable , with concrete examples covering deployment, monitoring, incident resolution, and lifecycle management. Deep hands-on expertise in Azure and Databricks , especially Azure Databricks, MLflow, and Unity Catalog. Experience with Azure ML is a strong advantage. Strong experience designing Databricks solutions, diagnosing workload bottlenecks, and producing clear technical designs and implementation proposals. Practical Python software-engineering skills for MLOps and platform work, including packaging, testing, APIs, and maintainable automation. Strong knowledge of containers and model-serving patterns. Experience with Kubernetes or AKS is beneficial but not required. Deep experience building CI/CD pipelines for models and data, beyond infrastructure and application deployment alone. Confidence with Terraform and Infrastructure as Code, particularly when integrating ML services into an existing Azure landing zone. Working knowledge of cloud networking and security, including IAM, encryption, secrets management, governance, and observability. Understanding of ML governance, lineage, auditability, and model documentation. A scientific or life-sciences background is not required. Working knowledge of GenAI technologies and patterns, including Azure OpenAI or Azure AI Foundry, RAG, vector search, evaluation, and guardrails. Ability to influence senior stakeholders, communicate technical strategy clearly, and collaborate effectively across AI, data, platform, security, and product teams. Strong mentoring skills and a pragmatic approach to helping engineering teams deliver independently. From the outset, you can independently: Deploy a model behind a production-ready API or managed serving endpoint. Configure MLflow experiment tracking and model registration. Build or improve a CI/CD pipeline for models and data. Assess an ML workload, identify bottlenecks, and produce a sound solution design and delivery proposal. Over time, you establish repeatable standards and platform capabilities that make ML delivery faster, safer, and more reliable across teams.
Benefits & conditions
Benefits What we offer A competitive compensation package A yearly education budget to steep your learning curve A yearly sport budget because a fit body leads to a fit mind A flexible working culture because your work-life balance matters to us A position that enables you to have an impact on 1’000s of people, and the whole company’s growth. An international, knowledgeable, and passionate team with a strong collaborative mindset
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps – What’s the deal behind it?
How to Become an AI Engineer
What Are Large Language Models?
MLOps And AI Driven Development