> Markdown version of [/jobs/ext/2391744-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2391744-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Reliability Engineer - **Company:** Intel Corporation - **Location:** Boxborough, MA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Big Data, Fault Tolerance, Failure Mode Effects Analysis, Python (Programming Language), Remote Access Service, SQL Databases, Data Analytics, Hardware Acceleration - **Published:** August 3, 2026 - **Apply:** https://dejobs.org/x/x/815330F3C8AC49269852682C9F632A60/job/ ## About the Role * BS/MS/PhD in EE/ME Reliability or related; and/or at least 4-6 yrs experience. * Experience authoring and owning reliability specs and requirement flow-down. * Strong RAS, FMEA, statistical reliability (Weibull, FIT) skills. * Experience with large-scale fleet telemetry and thermal/power redundancy., * AI cluster operations, data analytics (Python/SQL). ## Description Join us to help build the next generation of AI hardware solutions. You will be part of a highly skilled, agile team developing cutting-edge hardware for the AI domain, where we push the boundaries of what silicon can do for emerging AI workloads. With a startup-like culture, we move quickly and give engineers the opportunity to drive significant technical and business impact. We are continuously developing modern and effective working methods, including hands-on adoption of AI tools throughout the chip development flow. Mission: Define and own the pod-level reliability specifications that ensure the availability, resilience, and serviceability of a large-scale data center across hardware, thermal, and operational dimensions., * Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems. * Translate system/SLA requirements into pod and subsystem level reliability specs; flow requirements down to silicon, platform, and facilities teams. * Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates. * Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs. * Partner with facilities on pod power/cooling redundancy (N+1, 2N), thermal margins, and disaster-recovery readiness. * Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Staying Safe in the AI Future](https://www.wearedevelopers.com/videos/521-staying-safe-in-the-ai-future) - [Data Analytics with Microsoft Fabric: End-to-End Use Case with Data Agents](https://www.wearedevelopers.com/videos/1547-data-analytics-with-microsoft-fabric-end-to-end-use-case-with-data-agents) - [How Data is Shaping our Games](https://www.wearedevelopers.com/videos/176-how-data-is-shaping-our-games) ## Related Articles - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)