> Markdown version of [/jobs/ext/3117258-ml-ai-engineer](https://www.wearedevelopers.com/jobs/ext/3117258-ml-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML/AI Engineer - **Company:** Lloyds Banking Group - **Location:** Manchester, UK - **Salary:** £72,702.0 - £80,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Build Automation, BigQuery, Nvidia CUDA, Continuous Integration, Data Infrastructure, Monitoring of Systems, Python (Programming Language), Machine Learning, Release Management, Prometheus, Service Development Studio, Management of Software Versions, Istio, Grafana, Multi-Agent Systems, Kubernetes Helm Charts, Git, Integration Tests, Kubernetes, Low Latency, Google Cloud Functions, Machine Learning Operations, TensorRT, Software Version Control, Dynatrace, Docker - **Published:** September 28, 2026 - **Apply:** https://lbg.wd3.myworkdayjobs.com/LBG_Careers/job/Manchester/ML-AI-Engineer_148385-1 ## About the Role * Strong Python for automation, tooling, and service development. * Deep expertise in Kubernetes, Docker, Helm, operators, node-pool management, and autoscaling. * CI/CD expertise having hands-on experience with Harness (or similar) building multi-stage pipelines; experience with GitOps, artefact repositories, and environment promotion. * Practical experience with CUDA, TensorRT, Triton, TorchServe, and GPU scheduling/optimisation. * Proficiency in Prometheus, Grafana, Dynatrace defining SLIs/SLOs and alert thresholds for ML systems. * Experience operating MLflow (or equivalent) for experiment tracking, model bundling, and deployments. * Expert use of Git, branching models, protected merges, and code-review workflows. It would be great if you had any of the following… * Experience with GCP (e.g., GKE, Cloud Run, Pub/Sub, BigQuery) and Vertex AI (Endpoints, Pipelines, Model Monitoring, Feature Store). * Hooks for prompt/version management, offline/online evaluation, and human-in-the-loop workflows (e.g., RLHF) to enable continuous improvement. * Familiarity with Model Context Protocol (MCP) for tool interoperability, plus Google ADK, LangGraph/LangChain for agent orchestration and multi-agent patterns. * Ray, Kubeflow, or similar frameworks. * Experience embedding controls, audit evidence, and governance in regulated environments. * Experience with GPU efficiency, autoscaling strategies, and workload right-sizing. ## Description Exciting opportunity for a hands-on ML/AI Engineer to join our Data & AI Engineering team. You'll build, automate, and maintain scalable systems that support the full machine learning lifecycle. You will lead Kubernetes orchestration, CI/CD automation (including Harness), GPU optimisation, and large-scale model deployment, owning the path from code commit to reliable, monitored production services This is a unique opportunity to shape the future of AI by embedding fairness, transparency, and accountability at the heart of innovation. You'll join us at an exciting time as we move into the next phase of our transformation. We're looking for curious, passionate engineers who thrive on innovation and want to make a real impact., * Compose, build, and operate production-grade Kubernetes clusters for high-volume model inference and scheduled training jobs. * Configure autoscaling, resource quotas, GPU/CPU node pools, service mesh, Helm charts, and custom operators to meet reliability and efficiency targets. * Implement GitOps workflows for environment configuration and application releases. * Build CI/CD pipelines in Harness (or equivalent) to automate build, test, model packaging, and deployment across environments (dev / pre-prod / prod). * Enable progressive delivery (blue/green, canary) and rollback strategies, integrating quality gates, unit/integration tests, and model-evaluation checks. * Standardise pipelines for continuous training (CT) and continuous monitoring (CM) to keep models fresh and safe in production. * Deploy and tune GPU-backed inference services (e.g., A100), optimise CUDA environments, and leverage TensorRT where appropriate. * Operate scalable serving frameworks (NVIDIA Triton, TorchServe) with attention to latency, efficiency, resilience, and cost. * Implement end-to-end observability for models and pipelines: drift, data quality, fairness signals, latency, GPU utilisation, error budgets, and SLOs/SLIs via Prometheus, Grafana, and Dynatrace. * Establish actionable alerting and runbooks for on-call operations; drive incident reviews and reliability improvements. * Operate a model registry (e.g., MLflow) with experiment tracking, versioning, lineage, and environment-specific artefacts. * Enforce audit readiness: model cards, reproducible builds, provenance, and controlled promotion between stages ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)