> Markdown version of [/jobs/ext/1444694-sr-software-engineer-inference](https://www.wearedevelopers.com/jobs/ext/1444694-sr-software-engineer-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Software Engineer, Inference - **Company:** CoreWeave Europe - **Location:** London, UK - **Experience:** Expert - **Salary:** £321,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Nvidia CUDA, Continuous Integration, Distributed Systems, Python (Programming Language), Open Source Technology, Performance Tuning, Remote Direct Memory Access, Cloud Services, Prometheus, Software Engineering, Computer Networking Systems, High Performance Computing, Large Language Models, Grafana, Caching, Kubernetes, Information Technology, TensorRT - **Published:** July 26, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=3887fb6adf540dd5 ## About the Role * 3+ years of professional industry experience building distributed systems or high-performance cloud services. * Strong software development skills in Python or Go, alongside a deep familiarity with networked systems and performance optimization. * Hands-on experience with Kubernetes at production scale, including automated CI/CD and modern observability stacks (Prometheus, Grafana, OpenTelemetry). * Practical, working knowledge of inference internals: batching strategies, caching, mixed precision (BF16/FP8), and streaming token delivery. * Proven track record of improving tail latency (P95/P99) and service reliability through metrics-driven engineering work. * Experienced in defining and owning SLIs/SLOs, managing capacity planning, autoscaling policies, and driving post-incident remediation. * Bachelor's or Master's degree in Computer Science or a related technical field (or equivalent practical experience). Preferred: * Experience with low-level systems and high-performance computing components (e.g., C++ development, CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies). * Active open-source or production contributions to modern inference frameworks (such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe). * Experience leading multi-team technical initiatives or partnering directly with enterprise customers on mission-critical platform launches. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. * You love to own critical architectural spaces, breaking down complex multi-service engineering work into high-impact, achievable milestones. * You're curious about pioneering advanced model serving optimizations to maximize platform throughput and drive cost-per-token efficiencies. * You're an expert in metrics-driven engineering, systematically troubleshooting tail latency bottlenecks, and elevating coding and testing standards across a team. ## Description As a Senior Software Engineer, you will serve as an area owner responsible for leading technical designs, raising engineering standards, and delivering measurable improvements across multiple services. You will partner with product, orchestration, and hardware teams to evolve our Kubernetes-native inference platform while meeting strict P99 SLAs at scale. This role involves leading design reviews, decomposing multi-service work into clear milestones, and owning critical infrastructure areas such as request routing, adaptive scheduling, cost-per-token analytics, and GPU resource isolation. You will implement advanced optimizations-including micro-batch schedulers, speculative decoding, and KV-cache reuse-while evolving capacity planning, autoscaling policies, and mentoring engineers across the team. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)