AI Inference Engineer

Fuse Limited
London, UK
9 days ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Grid System Artificial Intelligence Nvidia CUDA Uptime Software Architecture Large Language Models Kubernetes Low Latency Slurm TensorRT

Requirements

partnering with the CUDA/GPU engineersMake the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plansAct as direct technical owner of inference performance and reliabilityWork closely with the CUDA and GPU engineering teams to integrate custom kernels and hardware performance work cleanly into the serving layerSet the standards, tooling and benchmarks this function will run on as it growsRequirements4+ years building or operating large-scale inference serving systems, or equivalent strong project/industry experienceDeep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)Strong systems thinking, able to reason about the full path from incoming request to served response across a large clusterComfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving systemA track record of making high-stakes architecture calls and owning the outcomeComfort operating without a playbook: this is a founding role shaping a new function, not joining an established oneBonus: Triton or custom ML inference/training frameworks; autoscaling or capacity planning for large-scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused computeBenefitsCompetitive salary and eligibility for equityBiannual bonus schemeFully expensed tech to match your needsPrivate health insuranceBreakfast and dinner allowance for office-based employeesAs we hire globally, benefits vary by location.Job SummaryID: 5E494FB5A4Department:Type: full time

About the company

DescriptionFuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.We’ve raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.We’re building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

2:22 min

Leveraging unique cultural backgrounds in engineering design

Ixchel Ruiz · LIVE

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

4:35 min

Defining service-level indicators based on user behavior

Maxim Schepelin Maxim Schepelin · World Congress 2026 Europe

Videos

See all

Related articles

See all