> Markdown version of [/jobs/ext/1323299-staff-data-center-operations-engineer](https://www.wearedevelopers.com/jobs/ext/1323299-staff-data-center-operations-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Data Center Operations Engineer - **Company:** Crusoe's Inc - **Location:** Denver, CO, United States - **Experience:** Expert - **Salary:** $150,000.0 - $170,000.0 - **Contract:** Permanent contract - **Skills:** Computing Platforms, Cloud Computing, Data Centers, Pattern Recognition, PCI Express, Runbook, Hardware Infrastructure - **Published:** July 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=6cfa7472841a3e60 ## About the Role Required * 7+ years in data center operations, field engineering, or OEM/ODM technical support with hands-on GPU infrastructure experience * Direct hands-on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms required * Deep familiarity with server platform architecture and OEM escalation and RMA processes * Experience leading or contributing to large-scale GPU cluster bring-ups including rack staging and production handoff * Demonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform-level issue resolution * Experience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teams * Strong written communication - comfortable producing escalation documentation, platform runbooks, and leadership reporting * Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%) Preferred * Direct experience with SuperMicro GPU platforms (B200, GB200, or newer); SuperMicro Certified Engineer credentials a plus * Familiarity with ASUS or Quanta server platforms and ODM engagement models * Experience with liquid-cooled GPU platforms and CDU integration * Familiarity with AMD Instinct GPU platforms (MI300X/MI350X/MI355X) * Prior experience at an AI cloud provider, hyperscaler, or GPU-first infrastructure operator * Experience contributing to technician certification programs or IC leveling standards within a DC ops organization ## Description Crusoe Cloud operates GPU infrastructure across six production sites globally, with a fleet that spans SuperMicro, HPE, and next-generation ODM platforms as we scale. We're looking for a Staff Data Center Operations Engineer to serve as the senior technical operations resource for the SiteOps org - based at our Denver headquarters, with cross-site scope and travel authority across our full portfolio. This role is the bridge between Crusoe's distributed site teams and our OEM and ODM hardware partners. You'll own platform-level escalations that exceed site-level capability, drive hardware decisions at the org level, and serve as SiteOps' technical presence at headquarters - visible to engineering, procurement, and leadership in a way that a field-based role cannot be. You'll be hands-on when the situation calls for it, traveling to sites for complex escalations, new platform bring-ups, and deployment support. But your primary leverage is organizational: building the technical standards, OEM relationships, and institutional knowledge that keeps Crusoe's GPU fleet reliable across every site., * Own Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution * Travel to sites as needed for complex platform issues, new hardware bring-ups, and deployment support * Identify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requests * Root-cause complex hardware issues - PCIe, BMC, thermal, fabric - and produce resolution documentation reusable across the SiteOps org * Hand off platform-level findings to the appropriate internal engineering teams with clear, well-documented escalation packages OEM & ODM Technical Partnership * Develop and maintain deep technical relationships with Crusoe's primary hardware partners - currently SuperMicro and HPE, with upcoming ODM's as growing platforms - at the engineering and field escalation level * Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements that benefit the full fleet * Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint * Support vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadership Platform Standards & Org Development * Own OEM platform technical knowledge at the SiteOps org level - escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites * Own the development and maintenance of platform-specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platforms * Design and deliver technical training for SiteOps technicians covering hardware architecture, platform-specific troubleshooting, and field procedures - both for new hire onboarding and ongoing skill development as the fleet and team evolve * Contribute to the technician certification program and technical leveling standards across the org * Support new site bring-up efforts providing platform readiness and deployment execution expertise HQ Presence & Cross-Functional Collaboration * Serve as SiteOps' senior technical representative at Crusoe HQ, participating in platform, engineering, and procurement discussions that affect site operations * Partner with engineering and procurement teams on sparing strategy, RMA lifecycle management, and OEM/ODM support contract structures * Provide operational input into next-generation GPU platform evaluations (GB300, VR200, and beyond) * Produce escalation reporting, platform health analysis, and operational insights for SiteOps leadership ## Related Videos - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [How to build a sovereign European AI compute infrastructure](https://www.wearedevelopers.com/videos/1102-how-to-build-a-sovereign-european-ai-compute-infrastructure) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [6 Emerging Technologies We’ll Learn About in 2025](https://www.wearedevelopers.com/magazine/381-6-emerging-technologies-we-ll-learn-about-in-2025) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data)