> Markdown version of [/jobs/ext/956375-principal-software-engineer-gpu-compute](https://www.wearedevelopers.com/jobs/ext/956375-principal-software-engineer-gpu-compute). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Engineer, GPU Compute - **Company:** Roblox - **Location:** San Mateo, CA, United States - **Experience:** Expert - **Salary:** $345,040.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Intelligent Platform Management Interface, Bootstrap (Software), Cloud Computing, Nvidia CUDA, Distributed Systems, Firmware, InfiniBand, Structured Programming, Kubernetes, Bare Metal - **Published:** June 5, 2026 - **Apply:** https://careers.roblox.com/jobs/7978492?gh_jid=7978492 ## About the Role * 10+ years of experience building and operating large-scale distributed systems and infrastructure. * Deep, hands-on GPU expertise at the machine management layer and above: GPU host provisioning, driver and firmware lifecycle, GPU health and reliability, and the realities of running accelerators in production. * A track record as an expert for compute, not just fleet management, with the scars to prove you have scaled GPU or accelerator infrastructure that other teams depend on. * Strong proficiency in Go or other well-structured programming languages. * Experience operating GPU and AI workloads in production, including familiarity with CUDA, GPU scheduling, and high-performance networking (NVLink, InfiniBand, RoCE). * Familiarity with Kubernetes for GPU workloads and with bare-metal concepts (firmware, BMC/IPMI/Redfish, OS imaging) is a strong plus. * A history of being the anchor expert that an organization relies on for its hardest GPU and compute problems, and the leadership to up-level the engineers around you. ## Description * Serve as the GPU technical leader for the Compute team, partnering across Kubernetes, Machine Bootstrap, Networking, and Cloud to drive GPU strategy end to end. * Own the GPU host lifecycle above raw fleet management: driver, firmware, and CUDA stack management, GPU health and telemetry, and remediation of GPU-specific failures (XID errors, ECC, thermal, NVLink and fabric faults). * Architect how GPU capacity is exposed to compute platforms, including scheduling, isolation, and integration with Kubernetes for GPU and AI workloads. * Drive GPU reliability and performance at fleet scale, defining the detection, diagnosis, and automated repair of unhealthy accelerators before they impact production. * Evaluate and onboard new GPU and AI accelerator platforms, networking topologies (NVLink, InfiniBand, RoCE), and multi-node training and inference patterns. * Establish the standards, tooling, and APIs that let other engineering teams consume GPU compute safely and efficiently, reducing toil and raising the bar for the org. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How we Build The Software of Tomorrow](https://www.wearedevelopers.com/magazine/120-how-we-build-the-software-of-tomorrow) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care)