> Markdown version of [/jobs/ext/2583379-sr-principal-design-quality-reliability-engineer-oci-data-center-infrastructure](https://www.wearedevelopers.com/jobs/ext/2583379-sr-principal-design-quality-reliability-engineer-oci-data-center-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Principal Design Quality & Reliability Engineer - OCI Data Center Infrastructure - **Company:** Oracle - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Centers, Failure Mode Effects Analysis, Oracle (Applications), Cloud Services, Data Driven Tests, Oracle Cloud Infrastructure - **Published:** August 4, 2026 - **Apply:** https://eeho.fa.us2.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1/requisitions/preview/341957 ## About the Role Experience in hyperscale or cloud data center infrastructure. Familiarity with AFR, IDR, MTBF, and reliability growth methodologies. Experience supporting GW-scale infrastructure deployments. Advanced degree in Engineering or related discipline preferred. Success Profile Recognized technical expert who influences through expertise rather than authority. Strong analytical and systems-thinking skills. Excellent communication and cross-functional collaboration. Passion for quality, reliability, and continuous improvement. Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. ## Description Oracle Cloud Infrastructure (OCI) is seeking a Sr Principal Design Quality & Reliability Engineer to serve as the technical authority for design and product quality across next-generation AI data center infrastructure. This individual contributor will define reliability strategies, influence engineering decisions, establish quality standards, and partner across Engineering, Product Engineering, Supply Chain, Construction, Operations, and suppliers to improve reliability at hyperscale. The role provides technical leadership without direct people management. Define and champion OCI's design quality and reliability strategy for critical infrastructure. Lead cross-functional design reviews focused on failure prevention, resiliency, maintainability, and lifecycle performance. Develop and apply reliability methodologies including FMEA, fault tree analysis, accelerated life testing, and design-for-reliability. Define qualification and acceptance criteria for critical infrastructure products and systems. Partner with suppliers to improve product quality, manufacturing readiness, and field reliability. Analyze field performance data, AFR, MTBF, warranty trends, and failure modes to identify systemic improvements. Recommend design, component, and supplier changes based on data-driven analysis. Influence product, technology, and supplier decisions through technical expertise and reliability insights. Develop KPI dashboards and benchmark OCI performance against industry best practices. Mentor engineers and provide technical leadership across organizations without direct management responsibility. ## Related Videos - [User 1st! Technology 2nd! Stop building AI nobody uses - start delivering real business outcomes](https://www.wearedevelopers.com/videos/100337-user-1st-technology-2nd-stop-building-ai-nobody-uses-start-delivering-real-business-outcomes) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [Leverage Cloud Computing Benefits with Serverless Multi-Cloud ML ](https://www.wearedevelopers.com/videos/78-leverage-cloud-computing-benefits-with-serverless-multi-cloud-ml) - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)