Hardware Engineer

STN, inc.
United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Computing Platforms Systems Engineering Intelligent Platform Management Interface Bash Shell BIOS Configuration Management Databases Computer Engineering Data Centers Linux Firmware Python (Programming Language) Systems Architecture
+3 more
Scripting Graphics Processing Unit (GPU) Information Technology

Job description

The Hardware Engineer owns hardware lifecycle for GPU and supporting infrastructure assets, including fleet health monitoring, RMA workflows, firmware management, and long-range capacity planning. The role is the technical owner of the physical compute platform., * Monitor GPU and server health including thermal, error rates, and component failures

  • Drive the RMA process with vendors (NVIDIA, Supermicro, HPE, and others) end-to-end
  • Manage firmware, BIOS, and BMC upgrade campaigns across the fleet
  • Develop hardware burn-in and acceptance test procedures, including NCCL and stress tests
  • Investigate hardware failures and produce vendor-grade root cause analyses
  • Maintain hardware inventory, asset records, and CMDB accuracy
  • Drive capacity planning across compute, storage, and networking
  • Coordinate with Procurement on spare parts strategy and stocking levels
  • Author hardware engineering runbooks and operational procedures
  • Support new platform bring-up, qualification, and reference architecture validation

Requirements

Do you have experience in System architecture?, * 5+ years in hardware engineering, systems engineering, or data center engineering

  • Deep knowledge of x86 server architecture, GPU systems, and modern storage
  • Hands-on experience with NVIDIA HGX, DGX, or hyperscale-class systems
  • Strong Linux fundamentals and scripting skills (Python, Bash)
  • Bachelor’s degree in computer science, electrical engineering, or related field, * Experience with NVIDIA Mission Control, Base Command Manager, or Bright Cluster Manager
  • Familiarity with IPMI, Redfish, and vendor management interfaces
  • Knowledge of liquid cooling and high-density power architectures
  • Experience operating fleets of 1,000+ GPUs

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 · LIVE

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · WWC 2021

1:44 min

Hardware availability through datacenter partners and SDKs

Christoph Lohrmann Christoph Lohrmann +1 · WWC Europe 2026

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all