> Markdown version of [/jobs/ext/569537-staff-site-reliability-engineer-observability-gcp](https://www.wearedevelopers.com/jobs/ext/569537-staff-site-reliability-engineer-observability-gcp). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - Observability GCP - **Company:** Okta, Inc. - **Location:** Bellevue, WA, United States - **Experience:** Experienced - **Salary:** $194,000.0 - $267,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Systems Engineering, Cloud Computing, Computer Programming, Software Debugging, DevOps, Distributed Systems, Domain Name System (DNS), Python (Programming Language), Linux Kernel, Reliability Engineering, Ruby, TCP/IP, Google Cloud, Load Balancing, Grafana, Kubernetes, Low Latency, Data Analytics, Software Coding, Terraform, Splunk - **Published:** June 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f9eca4314e26b78e ## About the Role Do you have experience in Tooling?, We are seeking a highly technical Observability Site Reliability Engineer with a specialty in Google Cloud, to own and expand our Observability ecosystem into GCP. In this role, you will move beyond simple monitoring to delivering a world class, comprehensive, scalable Observability Platform that enables our SRE teams and business partners. You will treat infrastructure as code-utilizing Terraform and strong coding proficiency in Go, Python, or Ruby-to automate the deployment of agents and collectors across complex distributed systems., GKE: Minimum 5+ Experience scaling and managing observability in a Google Cloud platform. Visualization: Expertise in creating intuitive, actionable Splunk or Grafana dashboards that correlate data across multiple sources.SRE Mindset: Minimum 3+ years of experience in an SRE, DevOps, or Systems Engineering role with a focus on high-availability systems. * Programming Proficiency: Strong coding skills in Python, Go for building internal tools and automating workflows. * Distributed Systems: Deep understanding of Linux internals, networking (TCP/IP, DNS, Load Balancing), and container orchestration (Kubernetes/GKE). * Problem Solving: A data-driven approach to debugging complex, cross-service performance bottlenecks. Bonus Skills (The "Nice-to-Haves") * Telemetry Standards: Hands-on experience with OpenTelemetry (OTel), Vector, or similar frameworks for instrumenting applications. * Grafana Loki: Experience in migrating Splunk to Grafana Loki Other Cloud Platforms: Experience managing observability native tools within AWS., * This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire. ## Description * Automated Infrastructure: Design, build, and maintain scalable observability infrastructure using tools like Terraform. * GCP Observabilty Engineering: Optimize the collection, processing, and storage of Observabilty data to ensure high reliability and low latency of our Splunk and Grafana services * Incident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "observability-driven development." * Automation: Eliminate "toil" by automating the deployment and scaling of observability agents and collectors. ## Related Videos - [Planet-Scale Dashboards](https://www.wearedevelopers.com/videos/1618-planet-scale-dashboards) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)