> Markdown version of [/jobs/ext/2704244-technical-staff-observability-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2704244-technical-staff-observability-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # technical Staff Observability Site Reliability Engineer - **Company:** Okta, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $194,000.0 - $267,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Systems Engineering, Cloud Computing, Computer Programming, Software Debugging, DevOps, Distributed Systems, Domain Name System (DNS), Python (Programming Language), Linux Kernel, Reliability Engineering, Ruby, TCP/IP, Load Balancing, Kubernetes, Low Latency, Data Analytics, Software Coding, Terraform, Splunk - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-site-reliability-engineer-splunk-okta-7793662 ## About the Role Log Management: Minimum 5+ Experience scaling and managing Splunk Cloud at scale (1000+ SVCs), including Workload Management (WLM) and HEC optimization. Visualization: Expertise in creating intuitive, actionable Splunk dashboards that correlate data across multiple sources. SRE Mindset: Minimum 5+ years of experience in an SRE, DevOps, or Systems Engineering role with a focus on high-availability systems. * Programming Proficiency: Strong coding skills in SPL, Go for building internal tools and automating workflows. * Distributed Systems: Deep understanding of Linux internals, networking (TCP/IP, DNS, Load Balancing), and container orchestration (Kubernetes/EKS). * Problem Solving: A data-driven approach to debugging complex, cross-service performance bottlenecks. Bonus Skills (The "Nice-to-Haves") * Telemetry Standards: Hands-on experience with OpenTelemetry (OTel), Vector, or similar frameworks for instrumenting applications. * Charge-back app: Experience in implementing Splunk charge-back app for usage reporting Cloud Platforms: Experience managing observability native tools within AWS or GCP. Additional requirements: * This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire. * This person must attend in person onboarding in our San Francisco office the first week of employment. ## Description We are seeking a highly technical Staff Observability Site Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to delivering a world class, comprehensive, scalable Observability Platform that enables our SRE teams and business partners. You will treat infrastructure as code-utilizing Terraform and strong coding proficiency in Go, Python, or Ruby-to automate the deployment of agents and collectors across complex distributed systems., * Automated Infrastructure: Design, build, and maintain scalable observability infrastructure using tools like Terraform. * Splunk Engineering: Optimize the collection, processing, and storage of log data to ensure high reliability and low latency of our Splunk services * Incident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "observability-driven development." * Automation: Eliminate "toil" by automating the deployment and scaling of observability agents and collectors. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)