> Markdown version of [/jobs/ext/3527077-machine-learning-engineer-remote](https://www.wearedevelopers.com/jobs/ext/3527077-machine-learning-engineer-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer (Remote) - **Company:** Ocho People - **Location:** Belfast, UK (Remote available) - **Experience:** Expert - **Salary:** £90,000.0 - £140,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, BigQuery, Python (Programming Language), Machine Learning, Azure Machine Learning, Large Language Models, Prompt Engineering, Information Technology, Virtual Agents, Data Pipelines - **Published:** October 1, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5904077609 ## About the Role * 5+ years in ML engineering, NLP or applied data science, with hands-on experience of LLM or agent-based systems * Practical experience building or operating evaluation frameworks, automated scoring or benchmark systems for ML or LLM outputs * Strong Python and experience building data pipelines for evaluation datasets * Solid understanding of NLP and modern LLM capabilities, including prompting techniques, agentic workflows and retrieval * A strong statistics foundation, including sampling, variance and judging whether a change in results is meaningful * Ability to design metrics that represent real-world quality, not just benchmark scores * Working experience with GCP (Vertex AI, BigQuery) or equivalent hands-on experience with another major cloud ML platform * A degree in Computer Science, Machine Learning, Statistics or a related field, or equivalent practical experience * Right to work in the UK Desirable / Nice to Have: * Experience with LLM-as-judge techniques, rubric design or human-in-the-loop evaluation programmes * Familiarity with agent architectures and the failure modes specific to multi-step agentic systems * Experience operating evaluation systems at scale in production ## Description This is a new role and the first dedicated evaluation hire, so you will define how the business measures the quality of its AI. The product is an AI agent, and every change to a model, prompt or piece of agent logic can quietly make it better or worse. Your job is to know which, before customers do. You will work closely with the engineers building agent capabilities and with the team's senior data scientist, and report into the Head of Data Science. The team is small, so breadth matters as much as depth, and you will regularly pick up work outside your core specialism. The people who thrive here are hands-on and curious about the business itself, not just the technology. You will be comfortable in the data, confident learning on the job, calm when priorities shift, and you will know when an evaluation is good enough to ship and when it needs more work. Cost is a real constraint in a start-up, so you will think about what each LLM call and pipeline costs as naturally as how accurate it is., * Design evaluation frameworks and metrics covering accuracy, safety, latency and cost across agent and LLM systems * Build benchmark eval sets that reflect real customer scenarios and edge cases * Develop automated scoring pipelines using rubric-based grading and LLM-as-judge techniques * Calibrate automated judges against human review so the team can trust the numbers * Stand up regression suites that catch quality drops from model, prompt or agent-logic changes before release * Create dashboards that track model and agent quality over time and across releases * Investigate failure modes in multi-step agent behaviour and prioritise fixes with the engineering team * Partner with the senior data scientist on deeper statistical analysis of results * Balance evaluation coverage against compute and API cost