Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview In this role you will define and own pod-level reliability specifications to ensure availability and resilience of large-scale data center hardware across compute, memory, storage, network, power, and cooling. You will translate system requirements into silicon and facility specs, lead failure analysis, and shape RAS features with a strong focus on actionable prevention. You work closely with cross-functional teams to drive reliability improvements and ensuring thriving AI workloads in a fast-moving, startup-like environment. This is a chance to impact system-level reliability at scale and contribute to cutting-edge AI hardware development. Compensation / Benefits * Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) across subsystems * Translate system/SLA requirements into pod and subsystem specs and flow requirements down to teams * Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions * Architect RAS features and graceful degradation/redundancy against pod specs * Partner with facilities on pod power/cooling redundancy, margins, and disaster-recovery readiness * Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec Tasks * Experience authoring and owning reliability specs and requirement flow-down * Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills * Experience with large-scale fleet telemetry and thermal/power redundancy * BS/MS/PhD in EE/ME Reliability or related; 4-6 yrs experience Key requirements * stock bonuses * health benefits * retirement plan * vacation * competitive pay
Requirements
Experteer Overview In this role you will define and own pod-level reliability specifications to ensure availability and resilience of large-scale data center hardware across compute, memory, storage, network, power, and cooling. You will translate system requirements into silicon and facility specs, lead failure analysis, and shape RAS features with a strong focus on actionable prevention. You work closely with cross-functional teams to drive reliability improvements and ensuring thriving AI workloads in a fast-moving, startup-like environment. This is a chance to impact system-level reliability at scale and contribute to cutting-edge AI hardware development. Compensation / Benefits * Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) across subsystems * Translate system/SLA requirements into pod and subsystem specs and flow requirements down to teams * Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective aaaaa on * Architect RAS features and graceful degradation/redundancy against pod specs * Partner with facilities on pod power/cooling redundancy, margins, and disaster-recovery readiness * Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec Tasks * Experience authoring and owning reliability specs and requirement flow-down * Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills * Experience with large-scale fleet telemetry and thermal/power redundancy * BS/MS/PhD in EE/ME Reliability or related; 4-6 yrs experience Key requirements * stock bonuses * health benefits * retirement plan * vacation * competitive pay
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Stephan Gillich - Bringing AI Everywhere
How to Become an AI Engineer
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Highest Paying Tech Companies for Developers