> Markdown version of [/jobs/ext/1303957-senior-ai-reliability-engineer-platform](https://www.wearedevelopers.com/jobs/ext/1303957-senior-ai-reliability-engineer-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Reliability Engineer (Platform) - **Company:** Flatiron Health - **Location:** UK - **Experience:** Expert - **Salary:** £76,147.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Automation of Tests, Cloud Engineering, Continuous Integration, Python (Programming Language), Machine Learning, Reliability Engineering, Azure Machine Learning, Software Safety, Runbook, Software Engineering, Management of Software Versions, Workflow Management Systems, Datadog, Data Logging, Large Language Models, Multi-Agent Systems, Apache Spark, Software Security, Model Validation, Gitlab-ci, Kubernetes, Machine Learning Operations, Splunk, Data Pipelines, Databricks - **Published:** July 17, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5804880164 ## About the Role We're looking for a Senior AI Reliability Engineer (Platform) to help us accomplish our mission to improve and extend lives by learning from the experience of every person with cancer. Are you ready to be the next changemaker in cancer care?, You're a senior technical practitioner with experience working across data science, machine learning, software engineering, platform engineering, or reliability engineering. You are comfortable operating in ambiguous spaces where the right answer is not always obvious, and you are motivated by turning emerging AI capabilities into production-ready systems that teams can actually trust. You understand that AI systems fail differently from traditional software. A model may not crash, but it may silently degrade, become less accurate, respond inconsistently, produce poor outputs, or create business risk in ways that are hard to detect without the right evaluation and observability patterns. This role is focused on that production behaviour and system health, not on pure model research or training. You likely have: * 5+ years of experience in platform engineering, SRE, machine learning, MLOps or a related technical field, with strong Python skills and experience building production-quality systems. * Experience designing experiments, evaluation frameworks, statistical analyses, and quality metrics for ML or AI systems, with familiarity in LLMs, RAG, AI agents, prompt evaluation, and model behaviour. * Strong understanding of AI reliability and observability, including logging, tracing, monitoring, drift detection, statistical analysis, uncertainty, alerting, and production system health. * Experience with modern cloud and ML infrastructure, including AWS, containers, Kubernetes, CI/CD, data pipelines, workflow orchestration, versioning, and distributed compute platforms. * Knowledge of agentic and multi-agent systems, including orchestration, state management, tool execution, governance, reliability, human-in-the-loop controls, and selecting the appropriate level of AI autonomy for a given problem. * Strong communication and collaboration skills, with the ability to explain AI behaviour and tradeoffs to technical and non-technical stakeholders and thrive in a fast-moving, ambiguous environment with a pragmatic, enablement-focused mindset. * Fluent in English. Optional * Experience with LLM evaluation, red-teaming, adversarial testing, AI safety, RAG evaluation, retrieval quality measurement, embedding drift, or AI observability and model monitoring. * Hands-on experience with observability and data/ML platforms such as Datadog, Splunk, OpenTelemetry, Databricks, Spark, Airflow, dbt, Ray, SageMaker, GitLab CI/CD, or similar technologies. * Experience working in healthcare, life sciences, or other regulated, privacy-sensitive environments. ## Description * Design, build, and continuously improve evaluation frameworks, benchmarks, and automated testing pipelines for AI, LLM-powered, and agentic workflows. * Define and monitor quality, reliability, safety, performance, and cost metrics for AI systems, including observability, drift detection, hallucination risk, retrieval quality, and end-to-end workflow behaviour. * Develop reliability engineering practices for AI-enabled systems, including SLOs, SLIs, monitoring, alerting, incident response, runbooks, and root-cause analysis of AI failure modes. * Design orchestration, governance, and guardrails for multi-agent AI systems, including agent coordination, permissions, auditability, human oversight, and secure deployment patterns. * Partner with platform, product, security, engineering, and data science teams to evaluate AI solutions, establish reusable standards, and guide build-vs-buy, model selection, and AI adoption decisions. * Support experimentation with emerging AI technologies while helping the organisation make pragmatic, scalable decisions in a rapidly evolving landscape, collaborating across global teams and participating in on-call rotations. ## Related Videos - [AI in High-Stakes Industries: Lessons Learned](https://www.wearedevelopers.com/videos/100253-ai-in-high-stakes-industries-lessons-learned) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Staying Safe in the AI Future](https://www.wearedevelopers.com/videos/521-staying-safe-in-the-ai-future) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)