Reliability Engineer

Intel Corporation
Boxborough, MA, United States
about 1 month ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Big Data Fault Tolerance Failure Mode Effects Analysis Python (Programming Language) Remote Access Service SQL Databases Data Analytics Hardware Acceleration

Job description

Join us to help build the next generation of AI hardware solutions. You will be part of a highly skilled, agile team developing cutting-edge hardware for the AI domain, where we push the boundaries of what silicon can do for emerging AI workloads. With a startup-like culture, we move quickly and give engineers the opportunity to drive significant technical and business impact.

We are continuously developing modern and effective working methods, including hands-on adoption of AI tools throughout the chip development flow.

Mission: Define and own the pod-level reliability specifications that ensure the availability, resilience, and serviceability of a large-scale data center across hardware, thermal, and operational dimensions., * Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems.

  • Translate system/SLA requirements into pod and subsystem level reliability specs; flow requirements down to silicon, platform, and facilities teams.
  • Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates.
  • Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs.
  • Partner with facilities on pod power/cooling redundancy (N+1, 2N), thermal margins, and disaster-recovery readiness.
  • Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec.

Requirements

  • BS/MS/PhD in EE/ME Reliability or related; and/or at least 4-6 yrs experience.
  • Experience authoring and owning reliability specs and requirement flow-down.
  • Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills.
  • Experience with large-scale fleet telemetry and thermal/power redundancy., * AI cluster operations, data analytics (Python/SQL).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

1:14 min

Evolution of distributed SQL database architectures

Wei Hu Wei Hu · World Congress 2024

1:32 min

Structuring platforms for new services and data analytics

Nevelina Aleksandrova · LIVE

7:08 min

Engineering practices for extreme platform reliability

Justin Kitagawa · Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

1:10 min

Introduction to Microsoft Fabric and data agents

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 · World Congress 2025

Videos

See all

Related articles

See all