> Markdown version of [/jobs/ext/2817816-mid-level-and-senior-ml-runtime-engineer](https://www.wearedevelopers.com/jobs/ext/2817816-mid-level-and-senior-ml-runtime-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Mid-Level and Senior ML Runtime Engineer - **Company:** Fractile - **Location:** Bristol, UK - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Software Engineering, Build Management, Information Technology - **Published:** September 10, 2026 - **Apply:** https://www.collegerecruiter.com/job/2840586713-mid-level-and-senior-ml-runtime-engineer ## About the Role We care most about depth of knowledge and a genuine interest in the problem space. You'll be a strong fit if you have: * Solid experience with ML inference at scale, including multi-user serving * A deep understanding of paged attention and inference engines such as vLLM * Familiarity with key components of the ML software ecosystem * Strong software engineering skills and an instinct for clean, maintainable systems, * Experience with Rust * Having built your own inference engine from scratch * A degree in Computer Science or a related field ## Description We're looking for a Senior ML Runtime Engineer to help us integrate Fractile's AI accelerators with the latest inference frameworks and build the runtime stack that makes them fly. You'll work on genuinely hard problems - KV cache management, scalable multi-user inference, and the internals of transformer model execution - alongside a collaborative team that values curiosity and rigor equally. This is a hybrid role, with offices in London and Bristol - your choice of base. What You'll Do * Integrate Fractile's AI acceleration hardware with leading inference engines including vLLM and SGLang * Research KV cache management technologies (including paged attention) and build proof-of-concept implementations tailored to our hardware * Work closely with the runtime team to design and build a scalable, bare-bones reference inference engine * Focus primarily on the transformer ML architecture * Share your expertise to help shape the direction of our runtime stack ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Answering the Million Dollar Question: Why did I Break Production?](https://www.wearedevelopers.com/videos/1171-answering-the-million-dollar-question-why-did-i-break-production) - [The Avengers Initiative (Practical Ethics for Software Engineers)](https://www.wearedevelopers.com/videos/2070-the-avengers-initiative-practical-ethics-for-software-engineers) - [Enabling intelligent logistics automation: home-grown Industrial IoT platform at Austrian Post](https://www.wearedevelopers.com/videos/2018-enabling-intelligent-logistics-automation-home-grown-industrial-iot-platform-at-austrian-post) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From Zero to Mobile Developer in 45 Minutes With SwiftUI](https://www.wearedevelopers.com/videos/183-from-zero-to-mobile-developer-in-45-minutes-with-swiftui) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)