> Markdown version of [/jobs/ext/2797144-senior-software-engineer-vllm](https://www.wearedevelopers.com/jobs/ext/2797144-senior-software-engineer-vllm). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer (vLLM) - **Company:** CommonAI CIC - **Location:** Cambridge, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Continuous Integration, Memory Management, Python (Programming Language), Open Source Technology, Performance Tuning, Software Safety, Software Engineering, Software Systems, System Programming, AI Infrastructure, Rust (Programming Language), Graphics Processing Unit (GPU), Pytorch, Large Language Models, HuggingFace, Hardware Acceleration, Machine Learning Operations, TensorRT - **Published:** September 8, 2026 - **Apply:** https://www.collegerecruiter.com/job/2840575727-senior-software-engineer-vllm ## About the Role We are seeking a Senior Software Engineer with a passion for open-source AI infrastructure to work on deploying, extending and optimising vLLM (and potentially other inference serving engines) to support our projects. You will play a crucial, high-impact role across both our key programmes, the Scaling Inference Lab () and the High Assurance programme. As an AI-first company, we strongly believe in collaborative engineering powered by the tools we are helping to build. This role places a major emphasis on using LLMs, AI coding assistants, and autonomous agents for software development., * Proven experience working as a Senior Software Engineer with deep expertise in both Python and low-level programming (e.g. C/C++, Rust, assembly or CUDA) * A history of direct contributions to vLLM or similar high-performance open-source ML/AI projects (e.g. PyTorch, Hugging Face TGI, TensorRT-LLM, Ray) * Strong understanding of LLM inference mechanics (e.g. KV caching, continuous batching, memory management, model quantisation) * Experience interacting with, and upstreaming code to, active open-source communities * Hands-on experience working on performance optimisation for hardware accelerators (GPUs, TPUs, CPU vector units or other accelerators) * A strong enthusiasm for using LLMs, coding assistants, and agents as core tools in your own software development process, * Experience working on software systems that operate in highly regulated or high-assurance environments (e.g. financial services) * An understanding of the latest AI safety research and active involvement in that community * Deep knowledge of modern MLOps practices, CI/CD, and large-scale deployments ## Description * Deploy, instrument and monitor open weight models served using vLLM * Implement new features within vLLM to support novel hardware architectures as part of the Scaling Inference Lab * Work with the Panopticon team to identify opportunities to extend vLLM to enhance accuracy, explainability, and accountability when using it to serve models in regulated environments * Actively collaborate with the open-source vLLM community to propose, review, and upstream core changes * Troubleshoot, profile, and optimise inference performance, focusing on latency, throughput, and hardware utilisation ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Three years of putting LLMs into Software - Lessons learned](https://www.wearedevelopers.com/videos/1508-three-years-of-putting-llms-into-software-lessons-learned) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)