> Markdown version of [/jobs/ext/1372749-ai-inference-engineer](https://www.wearedevelopers.com/jobs/ext/1372749-ai-inference-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Inference Engineer - **Company:** Fuse Limited - **Location:** London, UK (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Grid System, Artificial Intelligence, Nvidia CUDA, Uptime, Software Architecture, Autoscaling, Large Language Models, Build Management, Kubernetes, Low Latency, Slurm, TensorRT - **Published:** July 22, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=d361f6bf468995bd ## About the Role * 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience. * Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding). * Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster. * Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system. * A track record of making high-stakes architecture calls and owning the outcome. * Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one. Nice to Have * Experience with Triton or custom ML inference/training frameworks. * Experience with autoscaling or capacity planning for large-scale inference workloads. * Exposure to multi-tenant serving or SLA-driven infrastructure. * Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. * Familiarity with Kubernetes/Slurm for cluster orchestration. * Interest or experience in energy markets, grid systems, or sustainability-focused compute. ## Description + Define Fuse's inference serving strategy and architecture from first principles. + Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads. + Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers. + Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents). + Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans. + Act as a direct technical owner of inference performance and reliability. + Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer. + Set the standards, tooling, and benchmarks this function will run on as it grows. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)