System Administrator 4-IT
Oracle
United States
about 2 months ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on eeho.fa.us2.oraclecloud.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Bash Shell
Cloud Computing
Databases
Information Engineering
Linux
DevOps
Distributed Systems
Python (Programming Language)
Networking Basics
Oracle (Applications)
Reliability Engineering
+8 more
Cloud Services
Prometheus
Scripting
Kubernetes
Data Analytics
Oracle Cloud Infrastructure
Data Pipelines
Golang
Job description
- Own and improve service reliability, availability, and performance (SLO/SLA)
- Lead and participate in incident response, root cause analysis, and postmortems
- Develop and implement automation to reduce operational toil and improve efficiency
- Build and enhance monitoring, alerting, and observability frameworks
- Partner with OCI engineering teams to improve system design, scalability, and resilience
- Support production operations in a regulated, high-compliance environment
- Contribute to capacity planning and scaling strategies
- Partner with the SRE team to align AI Ops initiatives with reliability goals.
- Integrate AI-driven tools into observability platforms and incident management workflows.
- Provide recommendations for optimizing cloud resources and improving system resilience.
- Build dashboards and visualizations to present AI-driven insights to engineering and operations teams.
- Collaborate with data engineering teams to design data pipelines that aggregate and preprocess monitoring and log data from diverse cloud environments.
- Participate in on-call rotation for 24x7 service coverage
- Operations staff may be required to work on a rotating shift basis
Requirements
- 10+ years of experience in SRE, DevOps, or production engineering
- Strong hands-on experience with Oracle Cloud Infrastructure (OCI)
- Experience operating large-scale distributed systems in production
- Proficiency in one or more programming/scripting languages (Python, Go, Bash, etc.)
- Experience with monitoring, observability, and incident management practices
- Strong understanding of Linux systems and networking fundamentals
- Experience with CI/CD pipelines and infrastructure as code
- Knowledge/Experience with troubleshooting and managing databases (Oracle preferred)
- Preferred Qualifications:
- Experience supporting high-availability, customer-facing cloud services
- Background in regulated or government cloud environments
- Familiarity with FedRAMP, ILx, or similar compliance standards
- Experience with FedRAMP and 3PAO audit procedures and requirements
- Experience driving automation and reliability engineering best practices
- Experience with Shepherd or similar
- Experience with Kubernetes
- Experience with M&O stack: Graphana, Prometheus or similar
- Familiarity with construction & engineering industry desired
- Familiarity with SRE principles, including incident response, SLIs/SLOs, and resilience engineering.
- Proven track record in building automation solutions for cloud operations or DevOps processes.
- Eligibility Requirement:
- Must be a United States Citizen and currently reside in the United States
- Candidates who do not meet these requirements will not be considered, * Experience supporting high-availability, customer-facing cloud services
- Background in regulated or government cloud environments
- Familiarity with FedRAMP, ILx, or similar compliance standards
- Experience with FedRAMP and 3PAO audit procedures and requirements
- Experience driving automation and reliability engineering best practices
- Experience with Shepherd or similar
- Experience with Kubernetes
- Experience with M&O stack: Graphana, Prometheus or similar
- Familiarity with construction & engineering industry desired
- Familiarity with SRE principles, including incident response, SLIs/SLOs, and resilience engineering.
- Proven track record in building automation solutions for cloud operations or DevOps processes.
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on eeho.fa.us2.oraclecloud.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago
DC
Daniel Cranney
Stephan Gillich - Bringing AI Everywhere
almost 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
24 days ago
CH
Chris Heilmann
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
almost 2 years ago
AJ
Austin Joy
What Are The Top Skills Required For Azure Developers?
over 4 years ago