> Markdown version of [/jobs/ext/2733228-ai-infrastructure-systems-engineer-r-d-gpu-ai](https://www.wearedevelopers.com/jobs/ext/2733228-ai-infrastructure-systems-engineer-r-d-gpu-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Ai Infrastructure Systems Engineer (R&D / GPU / AI) - **Company:** Nebius Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $179,500.0 - $224,300.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, BIOS, Software Debugging, Linux, Firmware, Field-Programmable Gate Array (FPGA), Python (Programming Language), Network Troubleshooting, Linux Kernel, PCI Express, Performance Tuning, AI Infrastructure, CPLD, Network Switches, Graphics Processing Unit (GPU), Extensible Firmware Interface, High Performance Computing, Hardware Infrastructure, Network Server, Hardware Debugging - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/senior-ai-infrastructure-systems-engineer-rd-gpu-ai-nebius-8330415 ## About the Role * Minimum 5 years of hands-on systems engineering experience across Linux, server hardware, firmware, PCIe, and GPU or high-performance computing platforms. * Extensive Linux experience, particularly using Linux as an environment for hardware debugging, system-level troubleshooting, and platform investigation. * Knowledge of the Linux kernel and experience with kernel-level debugging or troubleshooting. * Strong knowledge of modern server architecture, particularly in high-performance, GPU-based environments. * Strong knowledge of NVIDIA GPU platforms and diagnostic tooling, including nvidia-smi, XID and SXID analysis, GSP and driver behaviour, NVLink, NVSwitch, Fabric Manager, DCGM diagnostics, and NCCL testing * Deep knowledge of PCIe, including topology, enumeration, root complexes, endpoints, switches, bridges, retimers, link speed and width, AER and DPC errors, completion timeouts, and link failures. Candidates should understand how protocol-level errors can relate to physical-layer or signal-integrity problems. * Experience benchmarking systems for performance, stability, power efficiency, thermal behaviour, and workload characteristics * In depth understanding of firmware interactions across BIOS/UEFI, BMC, CPLD/FPGA, GPU firmware, NIC firmware, and other platform components. * Experience with low-level hardware communication and debugging using I²C, SMBus, PMBus, register maps, bit masks, byte- and word-level data, and device datasheets. * Demonstrated ability to troubleshoot complex hardware, software, and networking issues. * Experience with deep problem investigation, root cause analysis, and performance optimization in cloud or high-performance computing environments. * Strong analytical and problem-solving skills with a performance-first mindset. * Scripting and automation using various languages (Python\Go), Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. ## Description The role involves forming technical hypotheses, designing targeted tests, analyzing system logs and low-level hardware data, and distinguishing between GPU, host, firmware, interconnect, thermal, and physical hardware failures. It requires strong ownership, curiosity, and the ability to turn ambiguous problems into evidence-based conclusions and repeatable troubleshooting procedures., This is a senior systems engineering role focused on diagnosing and resolving complex hardware and platform failures in high-performance AI infrastructure. This engineer will work across Linux, GPU servers, PCIe, firmware, power, and cooling to determine the true root cause of issues that may not have an existing troubleshooting procedure. ## Related Videos - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [How to Submit CFPs and Get into Public Speaking - Moran Weber](https://www.wearedevelopers.com/videos/2112-how-to-submit-cfps-and-get-into-public-speaking-moran-weber) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)