Machine Learning Engineer

Understanding Recruitment
Hertfordshire, UK
3 days ago
Apply on www.understandingrecruitment.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Cloud Computing Continuous Integration Software Debugging Monitoring of Systems Python (Programming Language) Machine Learning Natural Language Processing Software Engineering Retrieval-Augmented Generation Large Language Models
+10 more
Grafana Reliability of Systems Indexer Langfuse — LLM Observability and Analytics Platform Agentic-AI Fastapi AI Platforms Deployment Automation Machine Learning Operations Cloudwatch

Job description

Machine Learning Engineer

If you’re an ML Engineer who enjoys shipping and running AI systems in production more than spending your days experimenting with models, this could be a very good fit.

What’s in it for you?

  • Work on AI and LLM systems that are genuinely running in production
  • Own problems across development, infrastructure and deployment, rather than being boxed into one area
  • Build with modern GenAI technologies including RAG, agentic AI and LLMs
  • Significant exposure to AWS architecture, MLOps, CI/CD and observability
  • Freedom to improve how AI services are deployed, monitored and scaled
  • Opportunity to take increasing technical ownership and potentially step into a Senior/Lead role
  • Remote working with the option to spend time in the office

What you’ll be working on

  • Building and operating production AI/LLM services
  • Designing and scaling cloud infrastructure in AWS
  • Improving CI/CD, infrastructure-as-code and automated deployments
  • Developing and debugging Python services using tools such as FastAPI and Pydantic
  • Building production RAG pipelines, including embeddings, indexing, retrieval and reranking
  • Implementing monitoring, tracing and observability across AI services
  • Improving system reliability, performance, compute efficiency and cost
  • Owning technical problems from development and staging through to production

What we’re looking for

You’ll ideally have 4+ years of relevant engineering experience, although depth of experience matters more than an exact number.

The strongest fit will be someone with:

  • A background in ML Engineering, MLOps, Platform Engineering or Software Engineering
  • Strong Python development experience
  • Hands-on experience building and operating systems in AWS
  • Experience deploying and maintaining production ML or AI services
  • Good understanding of CI/CD, containers and infrastructure-as-code
  • Experience with monitoring and observability tools such as Grafana, CloudWatch, Langfuse or similar
  • Some practical exposure to LLMs, RAG, NLP or generative AI
  • The confidence to take ownership of production systems and help guide other engineers

Requirements

You’ll ideally have 4+ years of relevant engineering experience, although depth of experience matters more than an exact number.

The strongest fit will be someone with:

  • A background in ML Engineering, MLOps, Platform Engineering or Software Engineering
  • Strong Python development experience
  • Hands-on experience building and operating systems in AWS
  • Experience deploying and maintaining production ML or AI services
  • Good understanding of CI/CD, containers and infrastructure-as-code
  • Experience with monitoring and observability tools such as Grafana, CloudWatch, Langfuse or similar
  • Some practical exposure to LLMs, RAG, NLP or generative AI
  • The confidence to take ownership of production systems and help guide other engineers

Benefits & conditions

  • Work on AI and LLM systems that are genuinely running in production
  • Own problems across development, infrastructure and deployment, rather than being boxed into one area
  • Build with modern GenAI technologies including RAG, agentic AI and LLMs
  • Significant exposure to AWS architecture, MLOps, CI/CD and observability
  • Freedom to improve how AI services are deployed, monitored and scaled
  • Opportunity to take increasing technical ownership and potentially step into a Senior/Lead role
  • Remote working with the option to spend time in the office

What you’ll be working on

  • Building and operating production AI/LLM services
  • Designing and scaling cloud infrastructure in AWS
  • Improving CI/CD, infrastructure-as-code and automated deployments
  • Developing and debugging Python services using tools such as FastAPI and Pydantic
  • Building production RAG pipelines, including embeddings, indexing, retrieval and reranking
  • Implementing monitoring, tracing and observability across AI services
  • Improving system reliability, performance, compute efficiency and cost
  • Owning technical problems from development and staging through to production

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.understandingrecruitment.com
Prepare application

Good distractions

Loading talks and stories from around this role…