Systems Research Engineer (AI Infrastructure & Distributed Systems

European Tech Recruit
Edinburgh, UK
4 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence C++ (Programming Language) Cloud Computing Profiling Computer Programming Microprocessors Distributed Systems Memory Management Fault Tolerance Systems Theories Python (Programming Language) Machine Learning
+9 more
Rapid Prototyping Process Distributed Caching AI Infrastructure Graphics Processing Unit (GPU) Load Balancing Cloud Platform System Caching AI Platforms Information Technology

Job description

A pioneering research-focused engineering team is seeking Systems Research Engineers to help shape the next generation of AI-native infrastructure. This role sits at the intersection of systems research and large-scale engineering, focusing on distributed architectures that support the training, serving, and deployment of advanced AI models.

The rapid evolution of large-scale AI models is transforming how modern computing systems are designed and deployed. A highly advanced research-driven engineering group is building the next wave of infrastructure that powers intelligent systems at scale-redefining how models are trained, served, and optimised across distributed environments.

This role offers a unique blend of hands-on engineering and forward-looking research, ideal for engineers who want to push the boundaries of distributed systems and AI infrastructure while working on real-world, high-impact platforms.

What You’ll Work On

  • Build and experiment with distributed system components tailored for data-intensive and AI-driven workloads.
  • Design scalable infrastructure capable of operating across diverse hardware environments including CPUs, GPUs, and accelerators.
  • Develop high-performance model serving systems with a focus on efficiency, scalability, and resilience.
  • Analyse system behaviour using profiling tools to uncover performance bottlenecks and optimisation opportunities.
  • Improve memory usage, caching strategies, and scheduling efficiency in large-scale inference systems.
  • Create solutions that enable low-latency, multi-tenant AI services in distributed environments.
  • Explore and prototype new approaches to inference architecture and cluster-level orchestration.
  • Translate technical innovations into tangible outcomes, including internal adoption and external publications.
  • Work closely with global teams to shape long-term infrastructure direction and strategy.

Requirements

  • PhD in Computer Science, Electrical Engineering, or a related discipline.
  • Strong foundation in distributed systems and operating systems principles.
  • Understanding of machine learning infrastructure and large-scale model serving.
  • Experience with systems-level programming in C/C++.
  • Proficiency in Python for experimentation and rapid prototyping.
  • Familiarity with distributed algorithms and system design trade-offs.
  • Experience using performance analysis and profiling tools.
  • Ability to communicate complex ideas clearly and work effectively in collaborative environments.
  • Contributions to recognised systems or machine learning conferences.
  • Hands-on experience with load balancing, fault tolerance, or cluster scheduling.
  • Exposure to distributed caching, state management, or high-performance cloud systems.
  • Experience building or optimising large-scale AI or cloud infrastructure.

Benefits & conditions

  • Be part of a team shaping the infrastructure behind next-generation AI systems.
  • Work on problems that combine deep technical research with real-world deployment.
  • Gain exposure to cutting-edge architectures in distributed computing and AI.
  • Collaborate with globally recognised experts in systems and machine learning.
  • Opportunity to publish, innovate, and influence future technology directions.
  • Accelerate your career in one of the fastest-growing areas of technology.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

2:36 min

Choosing between managed AI platforms and custom governance

Péter Farkas Péter Farkas · Europe 2026 Virtual

2:03 min

Solving complex engineering challenges in artificial intelligence deployment

Nico Axtmann · World Congress 2022

2:33 min

Maintaining prompt structures for prefix caching

Douglas Reiser Douglas Reiser · Europe 2026 Virtual

Videos

See all

Related articles

See all