> Markdown version of [/jobs/ext/3281900-senior-machine-learning-engineer-large-systems](https://www.wearedevelopers.com/jobs/ext/3281900-senior-machine-learning-engineer-large-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Machine Learning Engineer (Large Systems) - **Company:** EngineersOfAI - **Location:** Cambridge, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Code Review, Nvidia CUDA, Distributed Computing Environment, Python (Programming Language), Machine Learning, Performance Tuning, Software Engineering, Pytorch, Large Language Models, Deep Learning, Kubernetes, Information Technology, Hardware Acceleration, Machine Learning Operations - **Published:** September 8, 2026 - **Apply:** https://www.collegerecruiter.com/job/2840616163-senior-machine-learning-engineer-large-systems ## About the Role * Bachelor/Master's/PhD or equivalent experience in Machine Learning, Computer Science, Maths, Data Science, or related field. * Proficiency in deep learning frameworks like PyTorch/JAX. * Strong Python or C++ software development skills. * Expertise in deep learning from model training to optimisation and evaluation. * Capable of designing, executing and reporting from ML experiments. * Developed deep understanding of performance bottlenecks and how to overcome them. * Ability to move quickly in a dynamic environment. * Enjoy cross-functional work collaborating with other teams. * Strong communicator - able to explain complex technical concepts to different audiences., + MLOps for Kubernetes-based clusters + Building production systems with large language models + Efficient computing based on low-precision arithmetic. * Experience writing C++/Triton/CUDA kernels for performance optimisation of ML models. * Experience in distributed training or inference of ML models across 64+ accelerators. * Familiarity with HPC systems and networking including Infini. ## Description As a Senior Machine Learning Engineer in the Applied AI team at Graphcore, you will contribute to advancing AI technology by developing and optimising AI models tailored to our specialised hardware. You will work on large-scale systems where performance is critical to the success of our projects. Working closely with the Software Development and Research teams, you will play a critical role in identifying opportunities to innovate and differentiate Graphcore's technology. We seek engineers with strong technical skills and an understanding of AI model implementation at scale, eager to make a tangible impact in this rapidly evolving field., * Implement latest machine learning models and optimise them for performance and accuracy, scaling to 1000s of accelerators. * Test and evaluate new internal software releases, provide feedback to software engineering teams, make necessary code fixes, and conduct code reviews. * Benchmark models and key ML techniques to identify performance bottlenecks and improve model efficiency. * Design and conduct experiments on novel AI methods, implement them and evaluate results. * Collaborate with Research, Software, and Product teams to define, build, and test Graphcore's next generation of AI hardware. * Engage with the AI community and keep in touch with the latest developments in AI. ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market)