> Markdown version of [/jobs/ext/1402838-site-reliability-engineer-lead](https://www.wearedevelopers.com/jobs/ext/1402838-site-reliability-engineer-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer, Lead - **Company:** Booz Allen Hamilton Inc. - **Location:** Chantilly, VA, United States - **Experience:** Expert - **Salary:** $99,000.0 - $225,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Confluence, JIRA, Cloud Computing, Cyber Security, Software Debugging, DevOps, Distributed Systems, Linux System Administration, Networking Basics, Reliability Engineering, Site Reliability Engineering Practices, Prometheus, Software Deployment, Data Logging, Scripting, Grafana, Reliability of Systems, Git, Containerization, Kubernetes, Rancher, Nessus, Cloudwatch, Terraform, AWS EKS, Docker, Elk Stack, Jenkins, Microservices - **Published:** July 23, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9049136/site-reliability-engineer-lead ## About the Role * 8+ years of experience with monitoring, logging, and observability platforms, such as Prometheus, Grafana, and ELK stack * 8+ years of experience with Linux systems administration and networking fundamentals within AWS * Experience with Python scripting and automation * Experience with Infrastructure as Code using Terraform and Terragrunt * Knowledge of Kubernetes administration, troubleshooting, and operations. * TS/SCI clearance with a polygraph * Bachelor's degree and 8+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering, or 12+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering in lieu of a degree * Ability to obtain a Security+ CE, SSCP, CCNA-Security, or GSEC Certification within 6 months of start date * Experience with deploying and managing OpenTelemetry . * Experience with AWS CloudWatch, AWS EKS, and related AWS services * Experience managing Kubernetes environments through Rancher * Experience implementing SRE practices such as SLOs, SLIs, error budgets, and incident management * Experience with Jenkins, Git, Docker, Kubernetes, Nessus, JIRA, and Confluence * Knowledge of distributed systems, microservices architectures, and containerized workloads * Knowledge of NIST 800-53 and NIST-190 * Master's degree in a relevant field * Security+ CE, SSCP, CCNA-Security, or GSEC Certification ## Description As a Lead Site Reliability Engineer (SRE) on our team, you'll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and implementation of observability, automation, incident response, and operational best practices across cloud and air-gapped environments, while partnering closely with D ev O ps , infrastructure, and security teams to improve system resilience and reduce operational risk. The Lead SRE also drives root cause analysis, capacity planning, reliability standards, and continuous improvement initiatives to support highly available , efficient, and scalable services. This is your chance to further your skills in cloud infrastructure and technologies while continuing to grow your SRE experience. You'll build and support a reliable site for the environment in order to meet the development and maintenance requirements of systems and platforms. Work with the development and operation teams to evaluate the health, stability and reliability of systems and platforms. Design and develop technical tools to debug problems that occur in the deployment of applications, within specific platforms and systems. Join our efforts to strengthen our security posture and safeguard national interests. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023)