Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid) |

Pure Resourcing Solutions
Cambridge, UK
4 days ago
Apply on www.reed.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
£90,000.0 - £120,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Apache HTTP Server Spreadsheets Profiling Nvidia CUDA Python (Programming Language) NumPy Open Source Technology Reliability Engineering Prometheus AI Infrastructure Pytorch
+5 more
Large Language Models Grafana Deep Learning Pandas Information Technology

Job description

Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid) | £90k-£120k Nobody quite knows where their compute budget is actually going until someone builds the model that tells them. That’s this role. My client is a Cambridge-based non-profit that exists to stop different parts of the AI world quietly rebuilding the same infrastructure. Rather than a startup, a big enterprise, a government department and a university lab each working out GPU efficiency from scratch, they pool the hard problems and the expertise needed to solve them, so everyone moves faster. It’s early days for the organisation but there’s serious momentum and serious financial backing behind it already. They’re hiring Performance Engineers at junior and senior level, to sit at the sharp end of that mission. Day to day You’d sit between the research and engineering teams, pulling real numbers off live training and inference runs rather than working from theory. From there, the job is building the models and calculators that turn those numbers into an actual answer: will this optimisation help, would a different accelerator be worth the spend, is this architecture change going to pay for itself. Those answers don’t stay internal either, they shape what gets bought and how systems get built, for the organisation itself and for everyone else in the membership relying on that judgement.

Requirements

  • A degree in computer science, mathematics, or something adjacent
  • A track record of building performance models or calculators (Python or spreadsheet-based) that actually forecast how a system will behave
  • Hands-on GPU/accelerator code optimisation, CUDA or similar
  • Genuine understanding of how LLMs and deep learning models run on real hardware, training versus inference, matrix multiplication, KV-caching, that level of detail
  • Comfortable in profiling tools like Nsight or PyTorch Profiler, and monitoring stacks like Prometheus and Grafana
  • Python for data work, Pandas and NumPy, plus general scripting

Nice to have rather than essential: a postgraduate degree and research background (publications welcome), real depth on inference serving frameworks like vLLM, a stats background, and any open source or research contributions., * Analytical Thinking (Solving)

  • Apache Web Server
  • Communication Skills (Key)
  • Hewlett Packard (HP) Products
  • Performance Engineering
  • Performance Management Roles (System)
  • Problem Solving (Process)
  • Python Programming (Beginner)
  • Reliability Engineering
  • Stakeholder Management (Business)

Benefits & conditions

It’s a rare early seat at something with genuine backing and genuine ambition, where the work you do gets acted on rather than filed away. Competitive salary and pension, hybrid from a Cambridge office, and real exposure to people across the wider AI and academic scene.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.reed.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · World Congress 2026 Europe

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · World Congress 2025

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all