> Markdown version of [/jobs/ext/171844-staff-data-center-operations-engineer-gpu-hardware-architecture](https://www.wearedevelopers.com/jobs/ext/171844-staff-data-center-operations-engineer-gpu-hardware-architecture). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Data Center Operations Engineer, GPU Hardware Architecture - **Company:** Crusoe's Inc - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $179,000.0 - $218,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Big Data, Computer Clusters, Computer Engineering, Data Centers, Datasheets, InfiniBand, Python (Programming Language), PCI Express, Tensorflow, Systems Architecture, Diagnostic Tools, Reliability of Systems, Data Analytics - **Published:** May 20, 2026 - **Apply:** https://www.dice.com/job-detail/92e6d2a6-b33b-44f2-87b2-d97e52c0ee03 ## About the Role * Silicon & Fabric Mastery: Expert-level knowledge of NVIDIA (Hopper/Blackwell/Rubin) and AMD (Instinct) architectures. Mastery of the physical and logical layers of NVLink, NVSwitch, and InfiniBand. * Infrastructure Bridge-Building: Ability to translate "Silicon Data Sheets" into "Mechanical Engineering Requirements." You can explain how a GPU's specific heat-load profile affects CDU sizing and secondary loop design. * Data-Driven Diagnostics: Proficient in Python, Go, or Bash to build telemetry and health-check tools (utilizing DCGM and ROCm). Experience using large datasets or basic ML frameworks to build "Smart Monitoring" that filters critical health signals from noise. * Operational Reliability Analysis: Experience using failure telemetry to inform site-level sparing requirements and field-service workflows. * Thermal Management: Deep understanding of the operational realities of Direct-to-Chip (D2C) cooling, including fluid dynamics, pressure-drop curves, and the lifecycle of dripless couplings., * 10+ years in Hardware Engineering, Systems Architecture, or Data Center Infrastructure. * The "Consultant" Mindset: Proven track record of educating and influencing cross-functional teams (specifically Engineering and Operations). * GPU Authority: You have managed or architected GPU clusters at scale (thousands of nodes) at a hyperscaler, a GPU-specialized cloud, or a major silicon vendor. Education: B.S. or M.S. in Electrical Engineering, Computer Engineering, or a related technical field. ## Description We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical authority on GPU platforms within the Data Center Engineering and Operations organization. Your mission is twofold: act as the primary technical consultant to our Data Center Engineering team to ensure future facilities are built for next-gen silicon, and provide the Operations team with the specialized tooling, SOPs, and predictive strategies needed to maintain peak cluster health., * For DC Engineering: You are the internal consultant. You translate upcoming GPU power/thermal roadmaps (NVIDIA/AMD) into design requirements for our next-generation facilities. * For Site Operations: You are the "Technical Enabler." You develop the diagnostic tools and technical SOPs that enable field technicians to resolve complex GPU issues with surgical accuracy. * For Sourcing: You are the "Technical Strategist." You define the technical sparing requirements and site-level inventory needs based on hardware failure telemetry., * Engineering Education & Design Support: Provide deep-dive technical guidance to the Data Center Engineering team on upcoming silicon (e.g., NVIDIA Blackwell/Rubin, AMD MI350/400). Ensure future facility designs for power, cooling, and rack-spacing are ready for 2000W+ per-chip densities. * Predictive Operations & Telemetry: Leverage AI/ML methodologies to analyze fleet-wide telemetry (power draws, thermal gradients, and error rates). You will lead the transition from reactive troubleshooting to predictive maintenance, identifying "pre-failure" patterns in HBM or NVLink components before they impact customer training runs. * Technical Sparing Architecture: Architect the site-level sparing strategy from a technical perspective. Use failure telemetry and MTBF data to define the "Critical Spares List" and stocking levels required at each site to meet cluster uptime targets, providing these requirements to Sourcing for execution. * Operational Tooling & SOPs: Build the "Operational Blueprint" for the field. Create precision SOPs for high-stakes GPU repairs (e.g., baseboard swaps, manifold maintenance) and develop diagnostic tooling that allows Site Ops to identify NVLink flapping, PCIe degradations, or thermal throttling. * Advanced Troubleshooting & RCA: Act as the Tier-3 escalation point for the most complex hardware failures in the production environment. Lead Root Cause Analysis (RCA) on systemic issues that span the boundary between hardware and facility environmental factors. * Silicon Roadmap Authority: Maintain a 24-month forward-looking view of NVIDIA and AMD architectures. Educate internal stakeholders on how transitions in HBM4, interconnect speeds, and liquid-cooling will impact Crusoe's physical infrastructure. * Vendor & VAR Technical Lead: Support the technical relationship with OEMs and VARs. Audit their hardware builds, review their technical bulletins, and ensure their hardware roadmaps align with Crusoe's operational and engineering standards. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Old tools, new tricks](https://www.wearedevelopers.com/videos/1916-old-tools-new-tricks) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [MCP doesn’t suck — your agent does](https://www.wearedevelopers.com/videos/100202-mcp-doesn-t-suck-your-agent-does) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)