System Administrator 4-IT

Oracle
United States
about 2 months ago
Apply on eeho.fa.us2.oraclecloud.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Bash Shell Cloud Computing Databases Information Engineering Linux DevOps Distributed Systems Python (Programming Language) Networking Basics Oracle (Applications) Reliability Engineering
+8 more
Cloud Services Prometheus Scripting Kubernetes Data Analytics Oracle Cloud Infrastructure Data Pipelines Golang

Job description

  • Own and improve service reliability, availability, and performance (SLO/SLA)
  • Lead and participate in incident response, root cause analysis, and postmortems
  • Develop and implement automation to reduce operational toil and improve efficiency
  • Build and enhance monitoring, alerting, and observability frameworks
  • Partner with OCI engineering teams to improve system design, scalability, and resilience
  • Support production operations in a regulated, high-compliance environment
  • Contribute to capacity planning and scaling strategies
  • Partner with the SRE team to align AI Ops initiatives with reliability goals.
  • Integrate AI-driven tools into observability platforms and incident management workflows.
  • Provide recommendations for optimizing cloud resources and improving system resilience.
  • Build dashboards and visualizations to present AI-driven insights to engineering and operations teams.
  • Collaborate with data engineering teams to design data pipelines that aggregate and preprocess monitoring and log data from diverse cloud environments.
  • Participate in on-call rotation for 24x7 service coverage
  • Operations staff may be required to work on a rotating shift basis

Requirements

  • 10+ years of experience in SRE, DevOps, or production engineering
  • Strong hands-on experience with Oracle Cloud Infrastructure (OCI)
  • Experience operating large-scale distributed systems in production
  • Proficiency in one or more programming/scripting languages (Python, Go, Bash, etc.)
  • Experience with monitoring, observability, and incident management practices
  • Strong understanding of Linux systems and networking fundamentals
  • Experience with CI/CD pipelines and infrastructure as code
  • Knowledge/Experience with troubleshooting and managing databases (Oracle preferred)
  • Preferred Qualifications:
  • Experience supporting high-availability, customer-facing cloud services
  • Background in regulated or government cloud environments
  • Familiarity with FedRAMP, ILx, or similar compliance standards
  • Experience with FedRAMP and 3PAO audit procedures and requirements
  • Experience driving automation and reliability engineering best practices
  • Experience with Shepherd or similar
  • Experience with Kubernetes
  • Experience with M&O stack: Graphana, Prometheus or similar
  • Familiarity with construction & engineering industry desired
  • Familiarity with SRE principles, including incident response, SLIs/SLOs, and resilience engineering.
  • Proven track record in building automation solutions for cloud operations or DevOps processes.
  • Eligibility Requirement:
  • Must be a United States Citizen and currently reside in the United States
  • Candidates who do not meet these requirements will not be considered, * Experience supporting high-availability, customer-facing cloud services
  • Background in regulated or government cloud environments
  • Familiarity with FedRAMP, ILx, or similar compliance standards
  • Experience with FedRAMP and 3PAO audit procedures and requirements
  • Experience driving automation and reliability engineering best practices
  • Experience with Shepherd or similar
  • Experience with Kubernetes
  • Experience with M&O stack: Graphana, Prometheus or similar
  • Familiarity with construction & engineering industry desired
  • Familiarity with SRE principles, including incident response, SLIs/SLOs, and resilience engineering.
  • Proven track record in building automation solutions for cloud operations or DevOps processes.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on eeho.fa.us2.oraclecloud.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all