> Markdown version of [/jobs/ext/2019473-staff-software-engineer-inference](https://www.wearedevelopers.com/jobs/ext/2019473-staff-software-engineer-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, Inference - **Company:** CoreWeave Europe - **Location:** London, UK - **Experience:** Expert - **Salary:** £81,006.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, C++ (Programming Language), Program Optimization, Nvidia CUDA, Distributed Systems, Python (Programming Language), Open Source Technology, Performance Tuning, Remote Direct Memory Access, Service Design, Computer Networking Systems, Cloud Platform System, System Availability, Large Language Models, Kubernetes, Information Technology, Machine Learning Operations, TensorRT, Software Coding - **Published:** August 11, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5834881302 ## About the Role * Minimum 8+ years of experience building large-scale distributed systems or cloud platforms. * Proven track record of leading cross-team or organization-level technical initiatives at scale. * Strong coding skills in Go, Python, or C++. * Deep expertise in Kubernetes at production scale, including orchestration, scheduling, and service design. * Strong understanding of networked systems, performance optimization, and distributed system design. * Hands-on engineering experience with inference systems, including batching/micro-batching strategies, caching, memory optimization, mixed precision (BF16/FP8), and streaming token delivery. * Demonstrated ability to systematically improve tail latency (P95/P99) and platform reliability through metrics-driven engineering. * Experienced in owning system-wide SLIs/SLOs, capacity planning, autoscaling strategies, and mentoring senior and mid-level engineers. * Bachelor's degree in Computer Science, Engineering, or a related technical field., * Direct open-source or production contributions to modern inference frameworks such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe. * Deep experience with GPU systems engineering and hardware performance optimization (CUDA, NCCL, RDMA, NUMA, or GPU interconnects). * Direct exposure to large-scale AI/ML infrastructure or hyperscale cloud environments. Wondering If You're a Good Fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. * You love to scale highly complex distributed architectures and mentor engineering cohorts to elevate technical standards across an organization. * You're curious about pioneering low-latency inference optimizations and finding innovative shortcuts to optimize cost-per-token performance. * You're an expert in metrics-driven engineering, troubleshooting micro-bottlenecks, and delivering stable cloud platforms under strict multi-tenant constraints. ## Description As a Staff Software Engineer on the Inference team, you will operate as a technical leader across multiple teams and services, driving architecture, performance, and reliability for CoreWeave's Kubernetes-native inference platform. You will define and lead complex, cross-cutting design initiatives spanning request routing, adaptive scheduling, GPU resource management, and cost-per-token optimization under strict P99 SLAs. This high-impact role requires you to implement advanced inference optimizations-such as speculative decoding and KV-cache reuse-while establishing performance benchmarking frameworks, guiding cross-functional alignment across infrastructure boundaries, and raising the bar for engineering rigor and observability practices. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)