Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Join us to help build the next generation of AI hardware solutions. You will be part of a highly skilled, agile team developing cutting-edge hardware for the AI domain, where we push the boundaries of what silicon can do for emerging AI workloads. With a startup-like culture, we move quickly and give engineers the opportunity to drive significant technical and business impact.
We are continuously developing modern and effective working methods, including hands-on adoption of AI tools throughout the chip development flow.
Mission: Define and own the pod-level reliability specifications that ensure the availability, resilience, and serviceability of a large-scale data center across hardware, thermal, and operational dimensions., * Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems.
- Translate system/SLA requirements into pod and subsystem level reliability specs; flow requirements down to silicon, platform, and facilities teams.
- Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates.
- Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs.
- Partner with facilities on pod power/cooling redundancy (N+1, 2N), thermal margins, and disaster-recovery readiness.
- Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec.
Requirements
- BS/MS/PhD in EE/ME Reliability or related; and/or at least 4-6 yrs experience.
- Experience authoring and owning reliability specs and requirement flow-down.
- Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills.
- Experience with large-scale fleet telemetry and thermal/power redundancy., * AI cluster operations, data analytics (Python/SQL).
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
How to Become an AI Engineer
Stephan Gillich - Bringing AI Everywhere
Highest Paying Tech Companies for Developers