ML Runtime Engineer

Fractile
Bristol, UK
1 day ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Programming Tools Device Drivers Firmware Software Engineering Build Management

Job description

About the Software organisation at Fractile

Developer Experience sits within the Software organisation at Fractile, which is responsible for developing a full software stack for our groundbreaking AI inference systems. That’s everything from ML compilers, device drivers and systems firmware, application level runtime and ecosystem integrations, ML and compute libraries, great developer tooling and a full portfolio of simulators, through to datacenter scale workload deployment solutions. At Fractile, we know that a fantastic software stack is a critical and central part of any AI inference solution and it sits at the heart of everything we’re doing.

About the team and role

The ML Runtime team is responsible for integrating Fractile’s AI accelerators with the latest inference frameworks and building the runtime stack that makes them fly. We work on genuinely hard problems - KV cache management, scalable multi-user inference, and the internals of transformer model execution - alongside a collaborative team that values curiosity and rigour equally.

As an ML Runtime Engineer you will integrate Fractile’s AI acceleration hardware with leading inference engines including vLLM and SGLang, research and build proof-of-concept KV cache management implementations tailored to our hardware, and work closely with the broader runtime team to design and build a scalable reference inference engine. You will focus primarily on the transformer ML architecture and share your expertise to help shape the direction of our runtime stack.

Requirements

You have solid experience with ML inference at scale, including multi-user serving, and a deep understanding of paged attention and inference engines such as vLLM. You are familiar with the key components of the ML software ecosystem and bring strong software engineering skills with an instinct for clean, maintainable systems. You care about depth of knowledge and have a genuine interest in the problem space - not just in shipping, but in understanding why things work the way they do., * Solid experience with ML inference at scale, including multi-user serving

  • Deep understanding of paged attention and inference engines such as vLLM
  • Familiarity with key components of the ML software ecosystem
  • Strong software engineering skills and an instinct for clean, maintainable systems, * Experience with Rust
  • Having built your own inference engine from scratch
  • A degree in Computer Science or a related field

Benefits & conditions

  • Competitive salary: A competitive salary reflective of your experience and the specialist nature of the role.
  • Equity & Ownership: meaningful equity so everyone shares in the value creation
  • Benefits: Private Medical, Dental and Vision, Contributory Pension, 25 Days holiday plus bank holidays and Life/Critical Illness Insurance.
  • Diverse & fun office: we believe the hardest problems get solved by the broadest range of minds. We are committed to Equal Employment Opportunity through attracting and retaining a diverse team and building an inclusive environment.

Fractile is seeking to increase the clock speed of global progress, one chip at a time. We’ve recently raised $220M from investors including Founders Fund and Accel and our most important work lies ahead. Join us!

About the company

Fractile was founded in 2022 on the bet that, eventually, the world’s most capable AI systems would be limited in their impact by the time taken to produce useful outputs. We bet everything on the logical conclusion: that the only way to truly unlock this latent value, to make speed viable at scale, was to radically re-invent the hardware that we run our frontier AI models on. Ever since, we have been building chips and systems that tackle this problem: how to efficiently generate output at thousands of tokens per second, while handling the complexity and capacity challenges of operating large models at very long contexts.

The workloads that push to the limits of the current frontier are already transformational; it is the technical and economic limits on inference speed that are constraining progress. The defining work of the 21st century will be marked by the engine of inference delivering immense and diffuse chains of intellectual inquiry, in drug discovery, in software engineering, in materials discovery, in any field where progress is driven by deep reasoning and intelligence to resolve complex problems.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:38 min

The convergence of mobile engineering and machine learning

Sasha Denisov Sasha Denisov ¡ World Congress 2026 Europe

5:11 min

Integrating flat external dependencies without build management tools

Jens Knipper Jens Knipper ¡ Europe 2026 Virtual

3:22 min

Comparing kernel modules with eBPF for system tracing

Ozan Sazak Ozan Sazak ¡ World Congress 2024

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac ¡ World Congress 2021

36 sec

Frustrations with the complexity of modern web development

Andrew Taylor ¡ Coffee With Developers

2:05 min

Layering abstraction APIs for visual computing in browsers

Desiree Santos Desiree Santos +1 ¡ World Congress 2024

Videos

See all

Related articles

See all