> Markdown version of [/jobs/ext/1793967-site-reliability-engineering-architect](https://www.wearedevelopers.com/jobs/ext/1793967-site-reliability-engineering-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineering Architect - **Company:** Oracle - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Big Data, Data Centers, Failure Mode Effects Analysis, Oracle (Applications), Reliability Engineering, Cloud Services, Enterprise Software Applications, Hardware Testing, Reliability of Systems, Oracle Cloud Infrastructure - **Published:** July 22, 2026 - **Apply:** https://eeho.fa.us2.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1/requisitions/preview/340746 ## About the Role * 10+ years of experience in reliability engineering, quality engineering, electrical design assurance, product quality, mission-critical power systems, data center infrastructure, industrial power, utilities, generation, transmission and distribution, or equivalent high-availability environments. * Strong systems-thinking capability, with demonstrated ability to evaluate electrical architectures under failure scenarios rather than only assessing individual components or equipment ratings. * Deep understanding of data center electrical distribution concepts, including medium-voltage and low-voltage distribution, substations, switchgear, UPS systems, generators, protection schemes, controls, grounding, power quality, redundancy, maintainability, and operational recovery. * Experience identifying failure modes, common-mode vulnerabilities, cascading risks, hidden dependencies, and blast-radius concerns across complex electrical systems. * Proven ability to assess availability impact and reliability tradeoffs using structured engineering methods such as FMEA, fault-tree analysis, reliability block diagrams, event analysis, root-cause analysis, or similar methodologies. * Relevant product quality and reliability experience across power infrastructure, electrical equipment, standardized electrical products, or repeatable data center design platforms. * Ability to influence design standards, supplier requirements, equipment qualification expectations, commissioning acceptance criteria, and operational readiness deliverables. * Strong cross-functional leadership skills, with experience partnering across design engineering, construction, commissioning, operations, vendors, and executive stakeholders. * Excellent communication and executive reporting skills, with the ability to translate complex technical reliability risk into clear decisions, priorities, and implementation plans. * Commitment to safety, compliance, disciplined engineering governance, and operational excellence in high-uptime environments. Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. ## Description At Oracle Cloud Infrastructure (OCI), we build the future of the cloud for Enterprises as a diverse team of fellow creators and inventors. We act with the speed and attitude of a start-up, with the scale and customer-focus of the leading enterprise software company in the world. Oracle Cloud Infrastructure is designing, standardizing, and operating mission-critical data center electrical infrastructure at extraordinary scale. We are hiring a Reliability & Quality Engineering leader to strengthen the design quality, system reliability, and availability of large-scale data center electrical distribution architectures. Early in the role, the priority will be evaluating the full data center electrical distribution system, identifying potential failure points, assessing how those risks could impact availability, and developing system-level resiliency concepts such as failure isolation, fault containment, and blast radius reduction. This role requires someone who can look beyond component-level reliability and understand how the overall electrical architecture behaves under credible failure scenarios. The ideal candidate will bring deep reliability and quality engineering capability across mission-critical power infrastructure, electrical equipment, or standardized data center design products. This person will partner closely with design engineering, construction, commissioning, operations, equipment suppliers, and executive stakeholders to convert reliability analysis into durable design standards, product improvements, risk mitigations, and measurable availability outcomes. What you'll do * Evaluate end-to-end data center electrical distribution architectures, including utility or behind-the-meter interfaces, substations, medium-voltage distribution, switchgear, UPS systems, generators, BESS where applicable, protection systems, controls, and downstream power delivery. * Identify design-level failure points, single points of failure, common-mode risks, hidden dependencies, protection coordination concerns, and failure modes that could materially impact availability. * Assess how electrical systems perform under credible failure scenarios, including equipment faults, transfer events, protection operations, control-system failures, degraded-mode operation, maintenance conditions, and abnormal grid or generation events. * Develop system-level resiliency concepts and design recommendations that improve fault isolation, recoverability, maintainability, failure containment, and blast radius reduction. * Partner with electrical design engineering, commissioning, operations, construction, supply chain, and equipment vendors to translate reliability findings into design standards, product requirements, test expectations, acceptance criteria, and corrective action plans. * Apply quality and reliability engineering methods such as FMEA, fault-tree analysis, reliability block diagrams, root-cause analysis, design-for-reliability reviews, lessons-learned integration, and field-performance trend analysis. * Evaluate product quality and reliability across power infrastructure, electrical equipment, standardized electrical products, and repeatable data center design platforms. * Recommend improvements to specifications, supplier qualification, equipment testing, design validation, manufacturing quality controls, inspection processes, and supplier feedback loops. * Create executive-ready reliability risk assessments that clearly communicate availability impact, technical trade-offs, mitigation options, residual risk, and prioritization across large-scale infrastructure programs. * Establish reliability KPIs, defect taxonomies, quality gates, issue review cadence, and design-readiness criteria that improve reliability before construction, commissioning, and operational handoff. * Champion safety, compliance, disciplined change management, and operational excellence while improving reliability at the architecture, product, and system levels. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [Leverage Cloud Computing Benefits with Serverless Multi-Cloud ML ](https://www.wearedevelopers.com/videos/78-leverage-cloud-computing-benefits-with-serverless-multi-cloud-ml) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)