Infrastructure and MLOps Engineer

Graphcore
Bristol, United Kingdom
9 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

Bristol, United Kingdom

Tech stack

Java
Artificial Intelligence
C++
Cloud Computing
Continuous Integration
Distributed Systems
Github
Python
Linux System Administration
Machine Learning
Cloud Services
Prometheus
Software Engineering
Software Systems
Datadog
High Performance Computing
Grafana
AI Platforms
Kubernetes
Infrastructure Automation Frameworks
Build Process
Terraform
Docker

Job description

Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible. Responsibilities and Duties

  • Develop, own, and maintain tools and services to support AI research and engineering teams

  • Deploy and maintain services with Kubernetes and Docker

  • Manage our Cloud Infrastructure using tools such as Terraform

Requirements

  • Knowledge of Python

  • Familiarity with cloud services (e.g. AWS)

  • Experience managing or developing in Linux environments

  • Understanding of CI/CD principles

  • Experience using Kubernetes (k8s)

  • Experience of one of the following:

  • maintaining machine learning applications.

  • deploying ML orchestration tools (e.g. NV Ray, KFP, SkyPilot).

  • managing ML accelerator hardware (e.g. DCGM).

Desirable

  • Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu)

  • Experience with GitHub Actions

  • Experience with modern observability tooling (e.g. Prometheus)

  • Experience with Grafana

  • Knowledge of Go/Java/C++ (or similar language)

About the company

As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence.

Apply for this position