Senior Systems HPC Engineer

Nebius
Amsterdam, Netherlands
28 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

C++ (Programming Language) Computer Clusters System Configuration Linux Network Interface Controllers InfiniBand Python (Programming Language) Kernel-Based Virtual Machine PCI Express Performance Tuning Quick EMUlator (QEMU) Software Engineering
+4 more
System Programming System Software Golang Programming Languages

Job description

We are looking for a Senior Systems HPC Engineer to play a key role in building our hyperscaler platform, working across its core components while analyzing and optimizing the performance of large-scale GPU clusters at the intersection of hardware and software.

You will operate across the full stack-from hardware and system software to networking (InfiniBand/RoCE), virtualization (KVM/QEMU), and distributed communication layers (e.g., MPI, NCCL).

In this role you will

  • Focus on understanding system behavior across multiple layers, identifying performance bottlenecks, and driving improvements that shape how our clusters are built, operated, tuned, and validated.
  • Investigate and troubleshoot performance issues of GPU cluster under real workloads (training and inference)
  • Evaluate and integrate new hardware, system configurations and tuning approaches through software stack
  • Support complex performance-related escalations from internal teams and customers
  • Work closely with infrastructure, software engineering and hardware vendor teams (e.g. NVIDIA, Mellanox, Intel)
  • Contribute to hardware and cluster qualification (acceptance), ensuring systems meet performance expectations

Requirements

  • 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming).
  • 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning).
  • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems.
  • Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python)., Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.

Benefits & conditions

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What’s it like to work at Nebius:

Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI

About the company

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.nl

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · WWC 2025

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · WWC Europe 2026

Videos

See all

Related articles

See all