> Markdown version of [/jobs/ext/2723101-staff-tdi-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2723101-staff-tdi-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff TDI Site Reliability Engineer - **Company:** Okta, Inc. - **Location:** Washington, DC, United States - **Experience:** Expert - **Salary:** $174,000.0 - $239,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Systems Engineering, Border Gateway Protocol, Software as a Service, Cloud Database, DevOps, Programming Tools, Monitoring of Systems, Internet Protocol Security (IP SEC), Internetworking, Python (Programming Language), Reliability Engineering, Software Engineering, Okta, Grafana, Amazon Virtual Private Cloud (VPC), Kubernetes, Infrastructure Automation Frameworks, AWS Fargate, Cloudwatch, Terraform, Splunk, Plan of Action and Milestones - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-tdi-site-reliability-engineer-okta-federal-okta-8777718 ## About the Role * 7+ years of experience as an SRE, DevOps Engineer, Cloud Automation Engineer, or Systems Engineer with a track record of delivering complex infrastructure projects at scale. * Experience with container orchestration and runtime environments, including EKS, ECS Fargate, and general container usage. * Proficient in infrastructure automation using Terraform and developing automation tools with Python, while leveraging secure software development practices. * Experience with monitoring tools, especially Splunk, CloudWatch, and the Grafana stack. * Experience with general networking concepts, such as BGP and IPsec management, and has leveraged AWS networking services, including VPCs, TGWs, and VPC endpoints. * Security Clearance: Active U.S. TS/SCI with polygraph. * The selected candidate may be subject to drug testing to the extent required by U.S. Government contracts. Additional requirements: * U.S. soil status - the employee must be on U.S. soil, which means the 50 states, the District of Columbia, or outlying areas of the United States, as defined in Federal Acquisition Regulation (FAR) 2.101 * U.S. Security Clearance status - the employee must be able to obtain and maintain a U.S. security clearance (Secret or Top Secret) to the extent required by U.S. Government contracts. ## Description Okta Federal, Inc. is looking for an experienced Staff TDI Site Reliability Engineer to help build, improve, and maintain our cloud platform services that help Okta support the most sensitive national security missions. The Site Reliability Engineering team delivers foundational infrastructure capabilities that enable corporate engineering teams to operate securely, reliably, and at scale. You'll play a key role in designing and implementing complex cloud-based engineering enablement systems, while ensuring compliance with strict government requirements. What you'll be doing * Operate and maintain enterprise grade solutions within air-gapped environments. * Build, run, and monitor development tools, pipelines, and infrastructure with a security-first mindset. * Operate autonomously within secure facilities. * Maintain SLOs/SLIs for workloads with no dependency on external monitoring or SaaS tooling. * Own runbooks and incident response procedures tailored to limited external escalation paths. * Participate in POA&M remediation and support annual/recurring Authority to Operate activities. * Support and run mission critical services depended on by product teams * Deliver excellent internal customer service and advocate for SRE and DevOps practices across teams. * Build and operate CI/CD pipelines that function without internet connectivity. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Reliable scalability: How Amazon.com scales on AWS](https://www.wearedevelopers.com/videos/983-reliable-scalability-how-amazon-com-scales-on-aws) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)