> Markdown version of [/jobs/ext/3603147-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/3603147-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** Understanding Recruitment - **Location:** Hertfordshire, UK (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Cloud Computing, Continuous Integration, Software Debugging, Monitoring of Systems, Python (Programming Language), Machine Learning, Natural Language Processing, Software Engineering, Retrieval-Augmented Generation, Large Language Models, Grafana, Reliability of Systems, Indexer, Langfuse — LLM Observability and Analytics Platform, Agentic-AI, Fastapi, AI Platforms, Deployment Automation, Machine Learning Operations, Cloudwatch - **Published:** October 7, 2026 - **Apply:** https://www.understandingrecruitment.com/job/machine-learning-engineer-9253/ ## About the Role You'll ideally have 4+ years of relevant engineering experience, although depth of experience matters more than an exact number. The strongest fit will be someone with: * A background in ML Engineering, MLOps, Platform Engineering or Software Engineering * Strong Python development experience * Hands-on experience building and operating systems in AWS * Experience deploying and maintaining production ML or AI services * Good understanding of CI/CD, containers and infrastructure-as-code * Experience with monitoring and observability tools such as Grafana, CloudWatch, Langfuse or similar * Some practical exposure to LLMs, RAG, NLP or generative AI * The confidence to take ownership of production systems and help guide other engineers ## Description Machine Learning Engineer If you're an ML Engineer who enjoys shipping and running AI systems in production more than spending your days experimenting with models, this could be a very good fit. What's in it for you? * Work on AI and LLM systems that are genuinely running in production * Own problems across development, infrastructure and deployment, rather than being boxed into one area * Build with modern GenAI technologies including RAG, agentic AI and LLMs * Significant exposure to AWS architecture, MLOps, CI/CD and observability * Freedom to improve how AI services are deployed, monitored and scaled * Opportunity to take increasing technical ownership and potentially step into a Senior/Lead role * Remote working with the option to spend time in the office What you'll be working on * Building and operating production AI/LLM services * Designing and scaling cloud infrastructure in AWS * Improving CI/CD, infrastructure-as-code and automated deployments * Developing and debugging Python services using tools such as FastAPI and Pydantic * Building production RAG pipelines, including embeddings, indexing, retrieval and reranking * Implementing monitoring, tracing and observability across AI services * Improving system reliability, performance, compute efficiency and cost * Owning technical problems from development and staging through to production What we're looking for You'll ideally have 4+ years of relevant engineering experience, although depth of experience matters more than an exact number. The strongest fit will be someone with: * A background in ML Engineering, MLOps, Platform Engineering or Software Engineering * Strong Python development experience * Hands-on experience building and operating systems in AWS * Experience deploying and maintaining production ML or AI services * Good understanding of CI/CD, containers and infrastructure-as-code * Experience with monitoring and observability tools such as Grafana, CloudWatch, Langfuse or similar * Some practical exposure to LLMs, RAG, NLP or generative AI * The confidence to take ownership of production systems and help guide other engineers