Compute Engineer, Deployment
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
This Compute Engineer will bring gigawatts of accelerators from first power-on to production. Facility availability to ready-for-service across thousands of racks per site, with a new data hall landing every few weeks.
They will make rack qualification faster than the fleet grows. Firmware baselines, burn-in, and cluster validation proven on every rack before a customer workload touches it, at a pace that never becomes the critical path.
They must be able to scale by tooling, not headcount. Deployed megawatts grow severalfold next year while the team stays near-flat, because anything done twice by hand becomes software.
Must own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.
Ability to qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.
Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.
Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.
Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.
Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.
Ability to travel 20-30% of the time to our Data Centers and Labs, as needed.
Requirements
This person has: brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.
Worked deep within Linux and out-of-band management: BMC, IPMI, and Redfish are daily tools for you, not occasional lookups.
Has automated hardware workflows in Python or Go rather than clicking through them, and the second time you do anything by hand you turn it into software.
Has worked physically in data halls, racking, cabling, and swapping components, and you’re just as effective acting as remote hands or directing them.
Be able to triage failures methodically across hardware, firmware, and software, isolating the fault to a component before reaching for a fix.
Willing to travel for turn-up windows when a new data hall comes online. Bonus: Kubernetes-based bare-metal provisioning. Accelerator platform bring up (NVIDIA, AMD, or custom). Burn-in and stress harness design. DCIM and inventory tooling.
About the company
Our client is a leading AI infrastructure company building and operating large-scale compute environments that power next-generation AI workloads. Their teams design, deploy, and operate high-performance data center infrastructure at massive scale, with a focus on speed, reliability, and operational excellence.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Highest Paying Tech Companies for Developers
Dev Digest 120 - Apple and peers
A Guide to Green Tech and Green IT Careers