Staff Site Reliability Engineer, Federal (TS/SCI)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Experteer Overview In this role you join Okta’s Federal SRE team to design, build, and operate reliable cloud services for government-focused workloads. You will lead reliability initiatives, drive automation, and collaborate with product, software, and security teams to meet FedRAMP/IL6 readiness. You’ll own incident response, SLI/SLO definitions, and continuous improvement of observability and platform tooling. This is a mission-driven opportunity to shape secure, scalable infrastructure for mission-critical applications. Compensation / Benefits * Design, build, and operate large-scale cloud infrastructure and production services * Participate in on-call rotation for highly available systems * Lead incident response and conduct post-incident reviews for systemic improvements * Define and improve SLIs, SLOs, and error budgets * Collaborate with engineering teams to enhance availability, scalability, and performance * Advance observability through metrics, logging, tracing, dashboards, and alerts * Develop automation and infrastructure code using Go, Python, Terraform * Promote CI/CD and GitOps practices to improve deployment safety * Mentor engineers and guide reliability engineering practices * Drive multi-team reliability initiatives and data-driven architectural influence Tasks * Extensive experience operating large-scale production services in AWS and/or GCP * Deep Kubernetes expertise in production environments * Proficiency with IaC (Terraform, Helm) * Strong software skills in Go and/or Python * Experience with distributed data stores (PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, etc.) * Solid understanding of cloud networking, TLS, DNS, load balancing, and service networking * Experience with observability platforms and production telemetry * Experience with or interest in AI-assisted engineering and automation * Proven incident response leadership and operational improvements * Security clearance readiness and ability to operate in regulated environments Key requirements * equity (where applicable) * bonus * health, dental and vision insurance * 401(k) * flexible spending account * paid leave
Requirements
- and alerts * Develop automation and infrastructure code using Go, Python, Terraform * Promote CI/CD and GitOps practices to improve deployment safety * Mentor engineers and guide reliability engineering practices * Drive multi-team reliability initiatives and data-driven architectural influence Tasks * Extensive experience operating large-scale production services in AWS and/or GCP * Deep Kubernetes expertise in production environments * Proficiency with IaC (Terraform, Helm) * Strong software skills in Go and/or Python * Experience with distributed data stores (PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, etc.) * Solid understanding of cloud networking, TLS, DNS, load balancing, and service networking * Experience with observability platforms and production telemetry * Experience with or interest in AI-assisted engineering and automation * Proven incident response leadership and operational improvements * Security clearance readiness and ability to operate in regulated aaaaaaaa lead Key requirements * equity (where applicable) * bonus * health, dental and vision insurance * 401(k) * flexible spending account * paid leave
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 120 - Apple and peers
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
Dev Digest 131 - AI'm not sure about OSS