Sr. Principal Design Quality & Reliability Engineer - OCI Data Center Infrastructure

Oracle
United States
about 1 month ago
Apply on eeho.fa.us2.oraclecloud.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Artificial Intelligence Data Centers Failure Mode Effects Analysis Oracle (Applications) Cloud Services Data Driven Tests Oracle Cloud Infrastructure

Job description

Oracle Cloud Infrastructure (OCI) is seeking a Sr Principal Design Quality & Reliability Engineer to serve as the technical authority for design and product quality across next-generation AI data center infrastructure. This individual contributor will define reliability strategies, influence engineering decisions, establish quality standards, and partner across Engineering, Product Engineering, Supply Chain, Construction, Operations, and suppliers to improve reliability at hyperscale. The role provides technical leadership without direct people management.

Define and champion OCI’s design quality and reliability strategy for critical infrastructure.

Lead cross-functional design reviews focused on failure prevention, resiliency, maintainability, and lifecycle performance.

Develop and apply reliability methodologies including FMEA, fault tree analysis, accelerated life testing, and design-for-reliability.

Define qualification and acceptance criteria for critical infrastructure products and systems.

Partner with suppliers to improve product quality, manufacturing readiness, and field reliability.

Analyze field performance data, AFR, MTBF, warranty trends, and failure modes to identify systemic improvements.

Recommend design, component, and supplier changes based on data-driven analysis.

Influence product, technology, and supplier decisions through technical expertise and reliability insights.

Develop KPI dashboards and benchmark OCI performance against industry best practices.

Mentor engineers and provide technical leadership across organizations without direct management responsibility.

Requirements

Experience in hyperscale or cloud data center infrastructure.

Familiarity with AFR, IDR, MTBF, and reliability growth methodologies.

Experience supporting GW-scale infrastructure deployments.

Advanced degree in Engineering or related discipline preferred.

Success Profile

Recognized technical expert who influences through expertise rather than authority.

Strong analytical and systems-thinking skills.

Excellent communication and cross-functional collaboration.

Passion for quality, reliability, and continuous improvement.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on eeho.fa.us2.oraclecloud.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

2:41 min

Transitioning artificial intelligence infrastructure into scalable commodity cloud services

juarezjunior juarezjunior · World Congress 2024

3:26 min

Parameterizing test functions with different input datasets

Florian Bruhin · World Congress 2021

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · World Congress 2024

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all