> Markdown version of [/jobs/ext/2245943-infrastructure-and-mlops-engineer](https://www.wearedevelopers.com/jobs/ext/2245943-infrastructure-and-mlops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure and MLOps Engineer - **Company:** Graphcore - **Location:** Cambridge, UK - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, C++ (Programming Language), Cloud Computing, Continuous Integration, Distributed Systems, Github, Python (Programming Language), Linux System Administration, Machine Learning, Cloud Services, Prometheus, Software Engineering, Software Systems, Datadog, High Performance Computing, Grafana, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Build Process, Machine Learning Operations, Terraform, Docker - **Published:** August 26, 2026 - **Apply:** https://www.collegerecruiter.com/job/2815168307-infrastructure-and-mlops-engineer ## About the Role * Knowledge of Python * Familiarity with cloud services (e.g., AWS) * Experience managing or developing in Linux environments * Understanding of CI/CD principles * Experience using Kubernetes (k8s) * Experience with one of the following: maintaining machine learning applications, deploying ML orchestration tools (e.g., NV Ray, KFP, SkyPilot), or managing ML accelerator hardware (e.g., DCGM) Desirable * Experience with Infrastructure as Code (IaC) tools (e.g., Terraform/OpenTofu) * Experience with GitHub Actions * Experience with modern observability tooling (e.g., Prometheus) * Experience with Grafana * Knowledge of Go/Java/C++ (or similar language) ## Description About Graphcore At Graphcore, we're building the future of AI compute. We're a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem. To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world. We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products, and the future of artificial intelligence. Job Summary Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team, enhancing the build, test, deployment, and productisation processes for our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems. The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible. Responsibilities And Duties * Develop, own, and maintain tools and services to support AI research and engineering teams * Deploy and maintain services with Kubernetes and Docker * Manage our Cloud Infrastructure using tools such as Terraform Candidate Profile Essential * Knowledge of Python * Familiarity with cloud services (e.g., AWS) * Experience managing or developing in Linux environments * Understanding of CI/CD principles * Experience using Kubernetes (k8s) * Experience with one of the following: maintaining machine learning applications, deploying ML orchestration tools (e.g., NV Ray, KFP, SkyPilot), or managing ML accelerator hardware (e.g., DCGM) Desirable * Experience with Infrastructure as Code (IaC) tools (e.g., Terraform/OpenTofu) * Experience with GitHub Actions * Experience with modern observability tooling (e.g., Prometheus) * Experience with Grafana * Knowledge of Go/Java/C++ (or similar language) Benefits In addition to a competitive salary, Graphcore offers flexible working, a generous annual leave policy, private medical insurance and a health cash plan, a dental plan, pension (matched up to 5%), life assurance and income protection. We have a generous parental leave policy and an employee assistance programme (including health, mental well-being and bereavement support). We offer a range of healthy food and snacks at our central Bristol office and have our own barista bar. We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal-opportunity process and encourage reasonable adjustments when required. ## Related Videos - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)