> Markdown version of [/jobs/ext/2450247-senior-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/2450247-senior-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Machine Learning Engineer - **Company:** CloudFlare - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Computer Vision, Distributed Systems, Python (Programming Language), Machine Learning, Open Source Technology, Regression Testing, Tensorflow, Graphics Processing Unit (GPU), Pytorch, Retrieval-Augmented Generation, Large Language Models, Deep Learning, Caching, Low Latency, ONNX (Open Neural Network Exchange) Format, Cloudflare, Machine Learning Operations, TensorRT, Serverless Computing - **Published:** August 27, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=708f9ea654ac2cbc ## About the Role * Experience building, optimizing, and operating machine learning models in production environments. * Strong proficiency with Python and modern ML frameworks such as PyTorch, TensorFlow, JAX, or equivalent. * Hands-on experience with inference optimization techniques for large-scale models, including quantization, batching, caching, compilation, and serving runtime tuning. * Experience with large-scale inference serving frameworks or runtimes such as SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp, or similar. * Familiarity with LLMs, speech models, vision models, embeddings, multimodal models, retrieval-augmented generation, or other modern deep learning architectures. * Experience optimizing models for GPUs or specialized accelerators. * Strong understanding of production ML concerns, including evaluation, monitoring, model regressions, rollout safety, and reliability. * Ability to work across ML and systems boundaries, including familiarity with distributed systems, networking, or serverless platforms. * Track record of leading complex technical projects and mentoring other engineers., * Experience contributing to open source ML tooling, model serving frameworks, or inference runtimes. ## Description At Cloudflare, we're not looking for people who wait for a polished roadmap; we're looking for the builders who see the cracks in the Internet that everyone else has simply learned to live with. We value candidates who have the instinct to spot a "normalized" problem and the AI-native curiosity to create a solution using the latest tools. Our culture is built on iteration, leveraging AI to ship faster today to make it better tomorrow, while ensuring that every improvement, no matter how small, is shared across the team to lift everyone up. If you're the type of person who values curiosity over bureaucracy, and that AI is a partner in solving tough problems to keep the Internet moving forward, you'll fit right in., You'll help define how machine learning models run across Cloudflare's global network, from frontier open LLMs and real-time voice models to customer-deployed models served on heterogeneous GPUs and next-generation accelerators. You'll work with systems engineers, product teams, hardware partners, and AI/ML engineers to bring models into production with low latency, strong reliability, and efficient resource use. This role combines applied ML, inference optimization, evaluation, and production engineering, with a focus on benchmarking models, improving serving performance, validating quality, and building tooling that helps Cloudflare and its customers ship AI applications at Internet scale., * Develop, optimize, and productionize machine learning models for Cloudflare's serverless inference platform, with a focus on performance, reliability, and model quality. * Build benchmarking and evaluation frameworks to measure latency, throughput, cost efficiency, and model behavior across LLMs, speech, vision, and other model families. * Improve inference performance through quantization, batching, caching, model compilation, runtime tuning, and accelerator-aware optimization. * Partner with systems engineers to integrate models into Cloudflare's distributed inference infrastructure across a heterogeneous fleet of GPUs and next-generation accelerators. * Drive improvements to model deployment workflows, including validation, rollout safety, observability, regression testing, and operational readiness. * Collaborate with product and engineering teams to translate customer requirements into scalable ML capabilities for Workers AI. * Mentor engineers, contribute to technical direction, and raise the quality bar for production ML engineering practices across the team. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How machine learning can help us tell fact from fiction](https://www.wearedevelopers.com/magazine/509-how-machine-learning-can-help-us-tell-fact-from-fiction) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)