> Markdown version of [/jobs/ext/1919092-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1919092-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Okta, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $165,000.0 - $225,600.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Amazon Web Services, Systems Engineering, Border Gateway Protocol, Cloud Computing Security, DevOps, Programming Tools, Disaster Recovery, Federal Information Processing Standards (FIPS), Github, Identity and Access Management, Internet Protocol Security (IP SEC), Python (Programming Language), Linux System Administration, Network Diagrams, Reliability Engineering, Runbook, Software Engineering, Data Logging, Grafana, Reliability of Systems, Amazon Virtual Private Cloud (VPC), Gitlab, Git, Containerization, Kubernetes, Cloudwatch, Terraform, Splunk - **Published:** August 4, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17809243?backUrl=%2Fcareer%2F17809243%2FSenior-Site-Reliability-Engineer-California-San-Francisco ## About the Role * Professional Experience & Scale: 5+ years of experience in SRE, DevOps, or Systems Engineering roles with a proven track record of delivering complex, large-scale infrastructure projects. * AWS Expertise & Centralized Governance: Expert in building and managing AWS multi-account environments (spanning hundreds of accounts), with deep proficiency in authentication, governance, and organization management (AWS Orgs, IAM, Identity Center, StackSets). * Automation & CI/CD Pipelines: Highly skilled in infrastructure as code (Terraform), writing secure automation tools in Python, and building Git-based CI/CD workflows (GitLab, GitHub Actions). * Containerization & Observability: Strong hands-on experience managing container orchestration environments (Kubernetes) and utilizing monitoring and logging tools (Splunk, CloudWatch, Grafana stack). Extra credit if you have experience in the following * Networking & AWS Infrastructure: Hands-on experience with general networking concepts (BGP and IPsec management) and leveraging core AWS networking services (VPCs, TGWs, and VPC endpoints). * System Administration: Solid foundational knowledge and experience in Linux system administration. * Security & Compliance: Proven experience operating within highly secure, regulated environments (e.g., FedRAMP), with a strong understanding of FIPS, STIGs, and data boundary implementations., * Federal Access & Eligibility: This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire. ## Description Reporting to the Manager, Site Reliability Engineering, this role will help build, improve, and maintain our cloud platform services by designing and implementing complex cloud-based engineering enablement systems. With a strong focus on automation, testing, and operational excellence, you will deliver foundational infrastructure capabilities that enable corporate engineering teams to operate securely, reliably, and at scale. What you'll be doing * Secure Cloud Infrastructure & Pipelines: Design, build, and modernize scalable cloud environments and development tools while strictly enforcing security policies and standards for regulated environments. * Cross-Functional Collaboration & Advocacy: Partner with software engineering teams to champion DevOps and SRE best practices, deliver excellent internal customer service, and actively contribute to Agile workflows (e.g., demos, architecture sessions). * Technical Documentation & Operations: Create and maintain comprehensive technical documentation, including network diagrams, runbooks, and disaster recovery procedures to ensure system reliability and knowledge sharing. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)