Staff Site Reliability Engineer, Federal (TS/SCI)

Okta, Inc.
Washington, DC, United States
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Cloud Computing Continuous Integration Distributed Data Store Domain Name System (DNS) Python (Programming Language) PostgreSQL MySQL Redis Reliability Engineering Software Tools
+10 more
Cloud Services Data Logging Transport Layer Security Load Balancing Okta Kubernetes Cassandra Data Analytics Terraform Golang

Job description

Experteer Overview In this role you join Okta’s Federal SRE team to design, build, and operate reliable cloud services for government-focused workloads. You will lead reliability initiatives, drive automation, and collaborate with product, software, and security teams to meet FedRAMP/IL6 readiness. You’ll own incident response, SLI/SLO definitions, and continuous improvement of observability and platform tooling. This is a mission-driven opportunity to shape secure, scalable infrastructure for mission-critical applications. Compensation / Benefits * Design, build, and operate large-scale cloud infrastructure and production services * Participate in on-call rotation for highly available systems * Lead incident response and conduct post-incident reviews for systemic improvements * Define and improve SLIs, SLOs, and error budgets * Collaborate with engineering teams to enhance availability, scalability, and performance * Advance observability through metrics, logging, tracing, dashboards, and alerts * Develop automation and infrastructure code using Go, Python, Terraform * Promote CI/CD and GitOps practices to improve deployment safety * Mentor engineers and guide reliability engineering practices * Drive multi-team reliability initiatives and data-driven architectural influence Tasks * Extensive experience operating large-scale production services in AWS and/or GCP * Deep Kubernetes expertise in production environments * Proficiency with IaC (Terraform, Helm) * Strong software skills in Go and/or Python * Experience with distributed data stores (PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, etc.) * Solid understanding of cloud networking, TLS, DNS, load balancing, and service networking * Experience with observability platforms and production telemetry * Experience with or interest in AI-assisted engineering and automation * Proven incident response leadership and operational improvements * Security clearance readiness and ability to operate in regulated environments Key requirements * equity (where applicable) * bonus * health, dental and vision insurance * 401(k) * flexible spending account * paid leave

Requirements

  • and alerts * Develop automation and infrastructure code using Go, Python, Terraform * Promote CI/CD and GitOps practices to improve deployment safety * Mentor engineers and guide reliability engineering practices * Drive multi-team reliability initiatives and data-driven architectural influence Tasks * Extensive experience operating large-scale production services in AWS and/or GCP * Deep Kubernetes expertise in production environments * Proficiency with IaC (Terraform, Helm) * Strong software skills in Go and/or Python * Experience with distributed data stores (PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, etc.) * Solid understanding of cloud networking, TLS, DNS, load balancing, and service networking * Experience with observability platforms and production telemetry * Experience with or interest in AI-assisted engineering and automation * Proven incident response leadership and operational improvements * Security clearance readiness and ability to operate in regulated aaaaaaaa lead Key requirements * equity (where applicable) * bonus * health, dental and vision insurance * 401(k) * flexible spending account * paid leave

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:41 min

Scale and diversity of software development teams

Bastian Heilemann Bastian Heilemann +1 · WWC 2025

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

2:33 min

Introduction to security advocacy and automation testing

Chris Heilmann +2 · LIVE

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · WWC 2022

1:48 min

Analyzing network packets with database protocol tools

Daniël van Eeden Daniël van Eeden · WWC Europe 2026

Videos

See all

Related articles

See all