Mid-Level and Senior ML Runtime Engineer

Fractile
Bristol, UK
18 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Working hours
Regular working hours

Tech stack

Artificial Intelligence Software Engineering Build Management Information Technology

Job description

We’re looking for a Senior ML Runtime Engineer to help us integrate Fractile’s AI accelerators with the latest inference frameworks and build the runtime stack that makes them fly. You’ll work on genuinely hard problems - KV cache management, scalable multi-user inference, and the internals of transformer model execution - alongside a collaborative team that values curiosity and rigor equally.

This is a hybrid role, with offices in London and Bristol - your choice of base.

What You’ll Do

  • Integrate Fractile’s AI acceleration hardware with leading inference engines including vLLM and SGLang
  • Research KV cache management technologies (including paged attention) and build proof-of-concept implementations tailored to our hardware
  • Work closely with the runtime team to design and build a scalable, bare-bones reference inference engine
  • Focus primarily on the transformer ML architecture
  • Share your expertise to help shape the direction of our runtime stack

Requirements

We care most about depth of knowledge and a genuine interest in the problem space. You’ll be a strong fit if you have:

  • Solid experience with ML inference at scale, including multi-user serving
  • A deep understanding of paged attention and inference engines such as vLLM
  • Familiarity with key components of the ML software ecosystem
  • Strong software engineering skills and an instinct for clean, maintainable systems, * Experience with Rust
  • Having built your own inference engine from scratch
  • A degree in Computer Science or a related field

Benefits & conditions

  • Work on one of the most technically ambitious projects in AI infrastructure
  • A small, expert team where your contributions are visible and valued
  • Hybrid working - split your time between home and our London or Bristol office
  • Competitive salary and equity
  • A culture that values learning, directness, and collaboration

About the company

We’re taking a revolutionary approach to computing - building AI acceleration hardware that runs the world’s largest language models 100× faster than existing systems. Our team works at the cutting edge of both hardware and software AI development, and we’re growing fast.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:54 min

Transitioning from software engineer to fractional engineering management

Lutz Heunkhen · Coffee With Developers

42 sec

Energy forecasts and resource demands of information technology

Marjolein Pordon · LIVE

1:22 min

Understanding software engineering as more than just coding

Lilia Gargouri Lilia Gargouri · World Congress 2026 Europe

5:11 min

Integrating flat external dependencies without build management tools

Jens Knipper Jens Knipper · Europe 2026 Virtual

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:46 min

Missing equipment retrieval processes for departing employees

Jasmin Azemović Jasmin Azemović · World Congress 2026 Europe

Videos

See all

Related articles

See all