> Markdown version of [/jobs/ext/2954078-mlops-engineer-llm-systems-serving-gpu-kernels-profiling](https://www.wearedevelopers.com/jobs/ext/2954078-mlops-engineer-llm-systems-serving-gpu-kernels-profiling). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) - **Company:** Weekday AI - **Location:** UK - **Experience:** Experienced - **Salary:** £104,416.0 - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Profiling, Nvidia CUDA, Software Debugging, Distributed Systems, Machine Learning, Pytorch, Large Language Models, Machine Learning Operations, TensorRT - **Published:** September 17, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5887356883 ## About the Role * 2+ years of hands-on professional experience in ML systems, ML infrastructure, model serving, or GPU and accelerator performance engineering. This is a hands-on systems role rather than an applied modelling or data science one. * Practical experience in at least one of the following, with more than one a strong plus: writing or optimizing custom GPU kernels (CUDA, Triton, Pallas); performance profiling and trace analysis (Kineto, torch.profiler, Nsight, XLA or JAX profiler); debugging distributed or accelerator-bound workloads; serving large language models at scale (vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, continuous batching). * Working production experience with JAX and/or PyTorch. Framework-level depth is a strong plus: custom operators, distributed training (FSDP, DDP, DeepSpeed, Megatron), or compiler and graph-level work. * Familiarity with modern accelerators such as A100, H100, B200 or TPU, and the ability to reason about throughput, latency and memory trade-offs. * Demonstrable career progression. * Ability to engage reliably for at least 40 hours/week during weekdays. * Strong written communication skills and the ability to explain complex technical decisions clearly. ## Description Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking MLOps Engineers with hands-on experience in large language model infrastructure across any of four areas: GPU kernel programming, performance profiling and trace analysis, debugging accelerated and distributed workloads, and high-throughput inference serving. This role involves AI model training and evaluation work, including writing and assessing MLOps and ML systems tasks and solutions to generate high-quality training data for frontier AI systems., * Design challenging, domain-relevant tasks across four areas, GPU kernels, performance profiling, debugging, and inference serving, and write accurate, well-structured solutions to them. * Guide research and engineering teams to close knowledge gaps and improve AI model performance on ML systems, training infrastructure, and framework-level topics. * Evaluate MLOps and ML systems tasks and solutions, and provide clear, written technical feedback that stands up to reviewer scrutiny. * Develop guidelines and detailed rubrics or evaluation frameworks covering kernel-level optimization, profiler output interpretation, distributed systems reasoning, and serving throughput and latency trade-offs. * Collaborate with other subject matter experts to keep training data consistent and accurate. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [MLOps - What’s the deal behind it?](https://www.wearedevelopers.com/videos/392-mlops-what-s-the-deal-behind-it) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering)