> Markdown version of [/jobs/ext/3424961-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3424961-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Oracle - **Location:** Austin, TX, United States - **Experience:** Expert - **Salary:** $81,100.0 - $187,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), JavaScript (Programming Language), Build Automation, Bash Shell, Unix, Cloud Computing, Configuration Management, Continuous Availability, Continuous Delivery, Continuous Integration, Dynamic Host Configuration Protocol, Linux, Distributed Systems, Domain Name System (DNS), Perl (Programming Language), Hypertext Transfer Protocols (HTTP), Python (Programming Language), Oracle (Applications), Performance Tuning, Reliability Engineering, Ruby, Software Deployment, Transmission Control Protocol (TCP), AI Infrastructure, Scripting, Cloud Platform System, Large Language Models, Kubernetes, Puppet, Terraform, Software Performance, Docker, Golang - **Published:** September 4, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9143728/senior-site-reliability-engineer ## About the Role * Linux and Unix operating systems * Docker, Kubernetes, and Terraform * Scripting languages such as Shell, Perl, Python, Java, and Go * Citizenship/location requirements - i.e. US Citizenship, U.S. Citizenship and possess and maintain TS/SCI w/Poly security clearance, reside in Austin, TX or Reston, VA. * Technology related bachelor's degree and/or equivalent work experience * A desire to learn and keep up with modern technologies * Proficient with writing services/task automation in Python, Bash, Ruby, Perl, JavaScript, or Java * Familiarity with core protocols (DNS, DHCP, HTTP, TCP) * Deep knowledge of Linux internals and host-based networking * Knowledge of Linux and/or Unix operating systems * Familiarity with configuration management solutions such as Chef, Puppet, etc * Experience with devising, managing, and extending monitoring solutions for large scale environments. * Knowledge of cloud computing concepts * Experience working in a mission-critical environment (Operations, Technical Support, NOC etc) * Proficient with communication skills (writing, organization, learning exchange) * Experience executing tasks under change management procedures * Experience resolving auto-cut and manual alarms following runbooks * A focus on customer satisfaction * Specific experience working with deployment of AI infrastructure to include clustered GPU's, LLM deployment and maintenance, and understanding of model integration integration for customer solutions. ## Description Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services. Design and develop designs, architectures, standards, and methods for large-scale distributed systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning You will provide cloud operations for Oracle National Security Realms. You'll be part of a dynamic team with a broad knowledge of how Oracle's cloud platform works. You'll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers. Note - this role is not a Monday to Friday core hours role - it will involve working a 24/7 shift rotation with on-call duties, including nights, weekends and public holidays. * Escalation points for junior Site Reliability Engineers during complex or high-impact incidents. * Manage and execute complex manual Change Management tickets, by working closely with the service teams to ensure safety and minimal disruption to services. * Support the on-boarding of new services and tools, ensuring they are operationally ready and properly integrated. * Provide mentorship and training to SREs, helping build team capability and confidence. * Create and maintain clear, useful documentation for operational processes and system support. * Identify areas of manual work and drive automation to reduce toil and improve efficiency. * Automate tasks to enable continuous delivery and ensure continuous availability with minimal human overhead * Recognize unsafe or inefficient practices and work with teams to design safer, more effective solutions. * Complete change requests to enable new functionality and maintain realm compliance * Ensure timely resolution of incidents, service requests, and change requests * Collaborate with global service and engineering teams * Define and drive change management, continuous integration, and deployment best practices * Help create and maintain real-world production architectures, scalability, and system design * Use a methodical approach to troubleshoot, large, complex, interconnected systems ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [WeAreDevelopers LIVE - Node and Package Security](https://www.wearedevelopers.com/videos/2138-wearedevelopers-live-node-and-package-security) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Coffee with Developers - Robby Russell](https://www.wearedevelopers.com/videos/917-coffee-with-developers-robby-russell) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)