Site Reliability Engineer 2

Oracle
UK
4 days ago
Apply on eeho.fa.us2.oraclecloud.com
Prepare application

Role details

Contract type
Contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) JavaScript (Programming Language) Agile Methodology Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Cloud Database Computer Programming Continuous Integration File Systems
+36 more
Distributed Systems Domain Name System (DNS) Issue Tracking Systems Python (Programming Language) Knowledge Management Network Troubleshooting Microsoft SQL Server MySQL Node.Js NoSQL Operational Data Store Oracle Databases Reliability Engineering Cloud Services Ansible Runbook Software Engineering Systems Integration Web Applications Enterprise Software Applications Load Balancing Bi Publisher Spring Cloud Sysadmin Delivery Pipeline Software Troubleshooting Git Oracle Service Cloud Kubernetes Infrastructure Automation Frameworks Information Technology Atlassian Tools Puppet Restful APIs Oracle Cloud Infrastructure Microservices

Job description

Site Reliability Engineer part of Oracle Analytics Service Excellence (OASE) UK GOV SRE team partnering with Oracle Analytics development teams to improve the reliability, availability, performance, operational support and maturity of Oracle Analytics Cloud services., * Perform SRE activities supporting Oracle Analytics customers, engineering teams, and release cycles in both pre-production and production environments.

  • Participate in a follow-the-sun model providing 24x7 operational support for Oracle Analytics services.
  • Respond to incidents, troubleshoot complex service issues, drive mitigation to completion, and contribute to root-cause analysis and post-incident actions.
  • Read, understand, and troubleshoot existing application and service code to diagnose complex production issues, identify root causes, and support safe remediation.
  • Develop deep expertise in Oracle Analytics services to prevent regressions, resolve customer issues effectively, and reduce recurring incidents.
  • Build, maintain, and improve operational tooling, automation, dashboards, monitoring, run-books, and knowledge-base documentation.
  • Build and maintain AI-assisted and automation tools as needed to improve SRE workflows, including incident investigation, monitoring, remediation, operational reporting, and knowledge management.
  • Develop scripts and services for monitoring, telemetry collection, capacity analysis, patching, remediation, and operational reporting.
  • Analyze service health, workload patterns, CPU, memory, query activity, and capacity trends to identify reliability and scaling risks.
  • Execute interim patches, hot fixes, upgrades, and maintenance activities with operational excellence.
  • Partner with Development, Support, Product Management, and other engineering teams to investigate and resolve service failures and outages.
  • Improve CI/CD processes, deployment practices, operational readiness, and release support-ability.
  • Share operational knowledge, and continuously improve team processes.
  • Follow established security, compliance, change-management, and operational procedures.

Requirements

The ideal candidate is a hands-on Site Reliability Engineer with strong analytical skills and the ability to read, understand, investigate, and safely troubleshoot existing application code. This role requires diagnosing complex service issues across distributed systems, automation, infrastructure, and enterprise applications-using logs, telemetry, code investigation, and operational data to drive issues from detection through mitigation and prevention., * BS or MS in Computer Science, Engineering, or equivalent practical experience.

  • Experience supporting cloud infrastructure, networking, applications, services, tools, and operational processes.
  • Strong understanding of networking and TCP/IP fundamentals, including DNS, HTTP/HTTPS, TLS, load balancing, and service connectivity.
  • Linux/Unix system-administration experience, including troubleshooting processes, memory, CPU, filesystem, and network issues.
  • Experience developing, operating, or supporting cloud services and large-scale distributed applications in production.
  • Demonstrated ability to troubleshoot complex technical issues methodically, including investigation of existing applications and code.
  • Experience creating and maintaining technical documentation, runbooks, knowledge articles, and operational guides.
  • Experience working in agile development and operational environments.
  • Strong written and verbal communication skills, including the ability to work effectively with remote global teams.
  • Ability to work independently, manage competing priorities, and participate in on-call, after-hours maintenance, and weekend support as needed., * Experience with Oracle Analytics Cloud, Oracle Analytics Server, OBIS, BI Publisher, Oracle Database, Autonomous Database, MySQL, SQL Server, or NoSQL technologies.
  • Two to four years of experience operating large-scale, customer-facing web applications or cloud services.
  • Experience with OCI, AWS, Azure, or GCP compute, storage, networking, monitoring, and operational tooling.
  • Programming and scripting experience with Python, Bash, JavaScript/Node.js, Ansible, and related technologies; Java experience is a plus.
  • Ability to read, understand, troubleshoot, and safely modify existing enterprise application code.
  • Familiarity with AI-assisted development tools, such as Codex and Claude Code, for software development, automation, investigation, and documentation.
  • Experience with CI/CD and infrastructure automation tools such as Ansible, Puppet, Chef, Git, and deployment pipelines.
  • Experience with cloud-native applications, containers, Kubernetes, microservices, and independently scalable services.
  • Experience with REST APIs, service integrations, and automation workflows.
  • Experience using Jira and Confluence for incident management, issue tracking, operational documentation, and collaboration.

About the company

OASE develops tools, technologies, processes, and data-driven operating practices that improve service uptime, reduce time to mitigation, and enable scalable cloud operations. The team builds and operates internal services, automation, dashboards, and reporting capabilities that support Oracle Analytics customers, engineering teams, partners, and business growth.

The ideal candidate enjoys working in an agile, customer-focused environment. This role is centered on improving uptime through proactive monitoring, incident response, code-level troubleshooting, automation through AI and tooling, capacity analysis, operational tooling, patching, remediation, and continuous service improvement.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on eeho.fa.us2.oraclecloud.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all