> Markdown version of [/jobs/ext/1572809-remote-site-reliability-developer-3-usc](https://www.wearedevelopers.com/jobs/ext/1572809-remote-site-reliability-developer-3-usc). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote - Site Reliability Developer 3 (USC) - **Company:** Oracle - **Location:** Denver, CO, United States - **Experience:** Expert - **Salary:** $79,100.0 - $158,200.0 - **Contract:** Permanent contract - **Skills:** Cerner, Bash Shell, Big Data, Encodings, Data Sharing, Linux, Distributed Systems, Apache Hadoop, Hadoop Distributed File System, Apache HBase, Python (Programming Language), Kerberos (Protocol), Oracle (Applications), Reliability Engineering, Ansible, Ruby, Apache Yarn, Core Data, Information Technology, Data Analytics, Apache Kafka, Terraform - **Published:** July 11, 2026 - **Apply:** https://dejobs.org/x/x/EE1D757410F8404DA5384687259D32C6/job/ ## About the Role * 4+ years operating large-scale, customer-facing distributed platforms * Deep experience with HDFS, YARN, HBase, Kafka, Storm, or similar systems * Strong background in Linux, networking, and distributed system troubleshooting * Infrastructure-as-Code using Ansible and Terraform * Scripting and automation using Python, Ruby, and Bash * Hands-on experience operating Kerberized environments * Proven ability to define and document technical architecture for complex systems * Demonstrated ownership of shared platforms with broad blast radius and multiple downstream consumers * Experience designing observability and capacity models for distributed platforms, * U.S. Citizenship and eligibility for a Federal Security Clearance * 5+ years of technical experience relevant to this position * Ability to communicate effectively and build rapport with team members * BS or MS in Computer Science, or equivalent ## Description This role provide support to core data platforms behind Oracle Health's Data & Analytics Platform. As a Senior Site Reliability Engineer (SRE), you will own shared, mission-critical systems used by multiple products and teams. You will work on the design and operation of large-scale, stateful distributed platforms, including Hadoop ecosystem components (HDFS, YARN, HBase) deployed on Oracle Big Data Service (BDS), Kafka, and Storm. These multi-tenant platforms are deployed and operated through Ansible- and Terraform-based automation and require strong architectural ownership to manage scale, change, and broad blast radius. What You'll Do Platform Ownership & Technical Leadership * Own the end-to-end reliability, scalability, and operability of shared data platforms * Define platform standards, architectural direction, and operational guardrails * Influence cross-team technical decisions and long-term platform strategy * Drive long-term platform evolution and influence reliability strategy across the data ecosystem Architecture & Design * Clearly articulate system behavior, dependencies, and failure modes * Make principled trade-offs between reliability, performance, cost, and complexity * Provide guidance and guardrails that enable downstream teams to use platforms safely and effectively Operations Engineering * Establish capacity models, scaling strategies, and operational best practices * Design platforms that behave predictably under load, failure, and change * Own platform lifecycle events: upgrades, expansions, decommissioning, and recovery Distributed Systems Expertise * Operate and evolve stateful distributed systems where data placement, replication, and recovery are critical * Reason about failure modes such as backpressure, rebalancing, region movement, replication lag, and rolling upgrades Security * Operate and maintain Kerberized platforms, including authentication, authorization, and secure service-to-service communication * Treat security as a first-class architectural concern Automation * Design and evolve an Ansible- and Terraform-driven automation framework * Treat automation as production software: versioned, reviewed, tested, and improved * Eliminate operational toil by encoding reliability and safety into the platform Incident Leadership & Prevention * Serve as the ultimate escalation point for complex or ambiguous incidents * Focus on eliminating entire classes of failure, not just resolving individual issues Representation * Represent SRE and platform engineering in high-visibility and sensitive forums * Communicate clearly with engineering leadership and partner teams Responsibilities Responsibilities The team operates within the Oracle Health Data & Analytics Platform, supporting one of Oracle Health's core products, HealtheIntent. We operate the big data and streaming infrastructure that enables downstream teams to deliver reliable customer-facing solutions at scale, while continuously improving operability and efficiency. ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Coroutine explained yet again 60 years later](https://www.wearedevelopers.com/videos/690-coroutine-explained-yet-again-60-years-later) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)