System Engineer (Compute Node)
Role details
Job location
Tech stack
Job description
The Compute Node team builds services for managing Virtual Machines running on GPU servers, integrating with disk management, virtual and InfiniBand networks. The team is also responsible for developing the Virtual Machine Scheduler that operates across clusters with thousands of servers and tens of thousands of GPUs, spanning multiple data centers in several regions.
Requirements
- Containerization & Orchestration
- Experience with Kubernetes not only for deployment but for development is highly valued.
- Virtualization
- KubeVirt and QEMU/KVM familiarity is a high plus.
- Linux & System-level Understanding
- Solid knowledge of internal OS architecture, performance considerations, process isolation and resource management.
- Knowledge of POSIX, sysfs, system calls, and file systems.
- Familiarity with server architecture, PCIe devices, NICs, and kernel drivers.
- Hardware Acceleration
- Experience or strong interest in working with GPUs, DPUs, or ARM architectures is desirable.
- Familiarity with the NVIDIA DOCA Software Framework will be an advantage.
- Go, C++ Proficiency
- Knowledge of at least one of these languages, ready to learn or work with the others as needed.
- Strong grasp of concurrency, debugging, and profiling techniques.
We conduct coding interviews as part of the process., Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
Benefits & conditions
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
What's it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI