> Markdown version of [/jobs/ext/1449671-staff-software-engineer-cloud-infrastructure](https://www.wearedevelopers.com/jobs/ext/1449671-staff-software-engineer-cloud-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer (Cloud Infrastructure) - **Company:** Crusoe's Inc - **Location:** San Francisco, CA, United States - **Salary:** $215,000.0 - $260,000.0 - **Contract:** Permanent contract - **Skills:** BIOS, Ubuntu (Operating System), CentOS, Command-Line Interface, Cloud Computing, Data Centers, Linux, Ethernet, Firmware, InfiniBand, Networking Hardware, Remote Direct Memory Access, Software Engineering, Diagnostic Tools, Graphics Processing Unit (GPU), Information Technology, Golang - **Published:** July 26, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=4ba414c0910febfd ## About the Role * Ability to code in Golang * Proven experience diagnosing and repairing high-density, rack-mounted compute hardware in production environments. * Deep understanding of GPU architectures and hands-on experience with GPU-based systems. * Experience supporting NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X series platforms. * Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE). * Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and testing. * Proficiency with GPU and system diagnostic tools such as NVIDIA DCGM and NVIDIA field diagnostic utilities. * Experience working with enterprise server hardware, power delivery, and cooling systems. * Strong analytical and problem-solving skills. * Excellent communication and collaboration skills. * Ability to work independently in a fast-paced data center or operations environment., * Technical certification or Associate's/Bachelor's degree in Electrical Engineering, Computer Science, or a related field or demonstrated experience. * Experience working directly with hardware vendors and escalations. * Background in large-scale GPU fleet operations or hyperscale data center environments. ## Description We are seeking a highly skilled and motivated GPU Fleet Operations Engineer to join Crusoe's Fleet Operations team. This role is focused on the advanced diagnosis, maintenance, and repair of high-performance GPU compute clusters, ensuring maximum uptime, reliability, and performance across our fleet., The ideal candidate will be hands-on with GPU rack-level troubleshooting and work closely with data center operations, engineering, and vendors to support cutting-edge infrastructure featuring the latest NVIDIA and AMD GPUs. This position plays a critical role in maintaining the health and scalability of Crusoe's rapidly growing GPU fleet., * Automate deep-level diagnosis and troubleshooting of hardware faults within GPU racks and high-density compute systems. * Develop software to troubleshoot and support GPU platforms including NVIDIA A100, H200, GB200, B200, B300 and AMD 350X / 355X. * Execute component-level diagnosis and remediation for failed or degraded hardware. * Partner with data center operations to manage and perform field-replaceable unit (FRU) repairs for GPUs, power supplies, cooling systems, interconnects, and networking hardware. * Conduct post-repair validation, burn-in testing, torch testing, and NVIDIA NCCL testing to ensure system stability and performance. * Implement and execute preventative maintenance procedures to improve fleet reliability and extend hardware lifespan. * Perform firmware and BIOS upgrades across the GPU fleet. * Maintain detailed documentation of maintenance activities, failures, and resolutions in ticketing and asset management systems. * Develop and update standard operating procedures (SOPs) for troubleshooting, repair, and validation workflows. * Collaborate with engineering, software, and data center operations teams to identify root causes of systemic failures and implement preventative solutions. * Participate in a rotating infrastructure on-call schedule (about one week every 4-6 weeks) with daytime coverage and handoff to the Europe team. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [Is Software Engineering Hard?](https://www.wearedevelopers.com/magazine/448-is-software-engineering-hard)