Reliability Engineer

Intel Corporation
Santa Clara, CA, United States
29 days ago
Apply on us.experteer.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Big Data Fault Tolerance Failure Mode Effects Analysis Hardware Design Remote Access Service Software Requirements Analysis Data Analytics

Job description

Experteer Overview In this role you will define and own pod-level reliability specifications to ensure availability and resilience of large-scale data center hardware across compute, memory, storage, network, power, and cooling. You will translate system requirements into silicon and facility specs, lead failure analysis, and shape RAS features with a strong focus on actionable prevention. You work closely with cross-functional teams to drive reliability improvements and ensuring thriving AI workloads in a fast-moving, startup-like environment. This is a chance to impact system-level reliability at scale and contribute to cutting-edge AI hardware development. Compensation / Benefits * Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) across subsystems * Translate system/SLA requirements into pod and subsystem specs and flow requirements down to teams * Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions * Architect RAS features and graceful degradation/redundancy against pod specs * Partner with facilities on pod power/cooling redundancy, margins, and disaster-recovery readiness * Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec Tasks * Experience authoring and owning reliability specs and requirement flow-down * Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills * Experience with large-scale fleet telemetry and thermal/power redundancy * BS/MS/PhD in EE/ME Reliability or related; 4-6 yrs experience Key requirements * stock bonuses * health benefits * retirement plan * vacation * competitive pay

Requirements

Experteer Overview In this role you will define and own pod-level reliability specifications to ensure availability and resilience of large-scale data center hardware across compute, memory, storage, network, power, and cooling. You will translate system requirements into silicon and facility specs, lead failure analysis, and shape RAS features with a strong focus on actionable prevention. You work closely with cross-functional teams to drive reliability improvements and ensuring thriving AI workloads in a fast-moving, startup-like environment. This is a chance to impact system-level reliability at scale and contribute to cutting-edge AI hardware development. Compensation / Benefits * Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) across subsystems * Translate system/SLA requirements into pod and subsystem specs and flow requirements down to teams * Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective aaaaa on * Architect RAS features and graceful degradation/redundancy against pod specs * Partner with facilities on pod power/cooling redundancy, margins, and disaster-recovery readiness * Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec Tasks * Experience authoring and owning reliability specs and requirement flow-down * Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills * Experience with large-scale fleet telemetry and thermal/power redundancy * BS/MS/PhD in EE/ME Reliability or related; 4-6 yrs experience Key requirements * stock bonuses * health benefits * retirement plan * vacation * competitive pay

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

1:32 min

Structuring platforms for new services and data analytics

Nevelina Aleksandrova · LIVE

1:05 min

Practical Byzantine Fault Tolerance in distributed computing systems

Jonan Scheffler · World Congress 2022

7:08 min

Engineering practices for extreme platform reliability

Justin Kitagawa · Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:25 min

Designing AI workflows for application fault tolerance

Daniel Oh Daniel Oh · World Congress 2026 Europe

Videos

See all

Related articles

See all