> Markdown version of [/jobs/ext/2767466-ai-inference-engineer](https://www.wearedevelopers.com/jobs/ext/2767466-ai-inference-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Inference Engineer - **Company:** Fuse Limited - **Location:** London, UK - **Contract:** Permanent contract - **Skills:** Grid System, Artificial Intelligence, Nvidia CUDA, Uptime, Software Architecture, Large Language Models, Kubernetes, Low Latency, Slurm, TensorRT - **Published:** September 7, 2026 - **Apply:** https://find.jobs/jobs-near-me/apply/ats-redirect/?id=2959819216-2 ## About the Role partnering with the CUDA/GPU engineersMake the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plansAct as direct technical owner of inference performance and reliabilityWork closely with the CUDA and GPU engineering teams to integrate custom kernels and hardware performance work cleanly into the serving layerSet the standards, tooling and benchmarks this function will run on as it growsRequirements4+ years building or operating large-scale inference serving systems, or equivalent strong project/industry experienceDeep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)Strong systems thinking, able to reason about the full path from incoming request to served response across a large clusterComfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving systemA track record of making high-stakes architecture calls and owning the outcomeComfort operating without a playbook: this is a founding role shaping a new function, not joining an established oneBonus: Triton or custom ML inference/training frameworks; autoscaling or capacity planning for large-scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused computeBenefitsCompetitive salary and eligibility for equityBiannual bonus schemeFully expensed tech to match your needsPrivate health insuranceBreakfast and dinner allowance for office-based employeesAs we hire globally, benefits vary by location.Job SummaryID: 5E494FB5A4Department:Type: full time ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)