Machine Learning Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Machine Learning Engineer
If you’re an ML Engineer who enjoys shipping and running AI systems in production more than spending your days experimenting with models, this could be a very good fit.
What’s in it for you?
- Work on AI and LLM systems that are genuinely running in production
- Own problems across development, infrastructure and deployment, rather than being boxed into one area
- Build with modern GenAI technologies including RAG, agentic AI and LLMs
- Significant exposure to AWS architecture, MLOps, CI/CD and observability
- Freedom to improve how AI services are deployed, monitored and scaled
- Opportunity to take increasing technical ownership and potentially step into a Senior/Lead role
- Remote working with the option to spend time in the office
What you’ll be working on
- Building and operating production AI/LLM services
- Designing and scaling cloud infrastructure in AWS
- Improving CI/CD, infrastructure-as-code and automated deployments
- Developing and debugging Python services using tools such as FastAPI and Pydantic
- Building production RAG pipelines, including embeddings, indexing, retrieval and reranking
- Implementing monitoring, tracing and observability across AI services
- Improving system reliability, performance, compute efficiency and cost
- Owning technical problems from development and staging through to production
What we’re looking for
You’ll ideally have 4+ years of relevant engineering experience, although depth of experience matters more than an exact number.
The strongest fit will be someone with:
- A background in ML Engineering, MLOps, Platform Engineering or Software Engineering
- Strong Python development experience
- Hands-on experience building and operating systems in AWS
- Experience deploying and maintaining production ML or AI services
- Good understanding of CI/CD, containers and infrastructure-as-code
- Experience with monitoring and observability tools such as Grafana, CloudWatch, Langfuse or similar
- Some practical exposure to LLMs, RAG, NLP or generative AI
- The confidence to take ownership of production systems and help guide other engineers
Requirements
You’ll ideally have 4+ years of relevant engineering experience, although depth of experience matters more than an exact number.
The strongest fit will be someone with:
- A background in ML Engineering, MLOps, Platform Engineering or Software Engineering
- Strong Python development experience
- Hands-on experience building and operating systems in AWS
- Experience deploying and maintaining production ML or AI services
- Good understanding of CI/CD, containers and infrastructure-as-code
- Experience with monitoring and observability tools such as Grafana, CloudWatch, Langfuse or similar
- Some practical exposure to LLMs, RAG, NLP or generative AI
- The confidence to take ownership of production systems and help guide other engineers
Benefits & conditions
- Work on AI and LLM systems that are genuinely running in production
- Own problems across development, infrastructure and deployment, rather than being boxed into one area
- Build with modern GenAI technologies including RAG, agentic AI and LLMs
- Significant exposure to AWS architecture, MLOps, CI/CD and observability
- Freedom to improve how AI services are deployed, monitored and scaled
- Opportunity to take increasing technical ownership and potentially step into a Senior/Lead role
- Remote working with the option to spend time in the office
What you’ll be working on
- Building and operating production AI/LLM services
- Designing and scaling cloud infrastructure in AWS
- Improving CI/CD, infrastructure-as-code and automated deployments
- Developing and debugging Python services using tools such as FastAPI and Pydantic
- Building production RAG pipelines, including embeddings, indexing, retrieval and reranking
- Implementing monitoring, tracing and observability across AI services
- Improving system reliability, performance, compute efficiency and cost
- Owning technical problems from development and staging through to production
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…