> Markdown version of [/jobs/ext/2603310-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2603310-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Reliability Engineer - **Company:** Intel Corporation - **Location:** Santa Clara, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Big Data, Fault Tolerance, Failure Mode Effects Analysis, Hardware Design, Remote Access Service, Software Requirements Analysis, Data Analytics - **Published:** August 5, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/reliability-engineer-santa-clara-ca-usa-58799663 ## About the Role Experteer Overview In this role you will define and own pod-level reliability specifications to ensure availability and resilience of large-scale data center hardware across compute, memory, storage, network, power, and cooling. You will translate system requirements into silicon and facility specs, lead failure analysis, and shape RAS features with a strong focus on actionable prevention. You work closely with cross-functional teams to drive reliability improvements and ensuring thriving AI workloads in a fast-moving, startup-like environment. This is a chance to impact system-level reliability at scale and contribute to cutting-edge AI hardware development. Compensation / Benefits * Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) across subsystems * Translate system/SLA requirements into pod and subsystem specs and flow requirements down to teams * Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective aaaaa on * Architect RAS features and graceful degradation/redundancy against pod specs * Partner with facilities on pod power/cooling redundancy, margins, and disaster-recovery readiness * Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec Tasks * Experience authoring and owning reliability specs and requirement flow-down * Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills * Experience with large-scale fleet telemetry and thermal/power redundancy * BS/MS/PhD in EE/ME Reliability or related; 4-6 yrs experience Key requirements * stock bonuses * health benefits * retirement plan * vacation * competitive pay ## Description Experteer Overview In this role you will define and own pod-level reliability specifications to ensure availability and resilience of large-scale data center hardware across compute, memory, storage, network, power, and cooling. You will translate system requirements into silicon and facility specs, lead failure analysis, and shape RAS features with a strong focus on actionable prevention. You work closely with cross-functional teams to drive reliability improvements and ensuring thriving AI workloads in a fast-moving, startup-like environment. This is a chance to impact system-level reliability at scale and contribute to cutting-edge AI hardware development. Compensation / Benefits * Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) across subsystems * Translate system/SLA requirements into pod and subsystem specs and flow requirements down to teams * Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions * Architect RAS features and graceful degradation/redundancy against pod specs * Partner with facilities on pod power/cooling redundancy, margins, and disaster-recovery readiness * Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec Tasks * Experience authoring and owning reliability specs and requirement flow-down * Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills * Experience with large-scale fleet telemetry and thermal/power redundancy * BS/MS/PhD in EE/ME Reliability or related; 4-6 yrs experience Key requirements * stock bonuses * health benefits * retirement plan * vacation * competitive pay ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [System Resilience: Surviving the Software Storm](https://www.wearedevelopers.com/videos/874-system-resilience-surviving-the-software-storm) - [Data Analytics with Microsoft Fabric: End-to-End Use Case with Data Agents](https://www.wearedevelopers.com/videos/1547-data-analytics-with-microsoft-fabric-end-to-end-use-case-with-data-agents) - [How Data is Shaping our Games](https://www.wearedevelopers.com/videos/176-how-data-is-shaping-our-games) ## Related Articles - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)