Lead Site Reliability Engineer

Quantum Science Solutions
Herndon, VA, United States
10 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Code Review Computer Networks DevOps Distributed Systems Python (Programming Language) Reliability Engineering Cloud Platform System Kubernetes Terraform Golang

Job description

Are You a Lead Site Reliability Engineer Ready to Drive Reliability, Automation, and Scalable Infrastructure? QSSHire is seeking an experienced The Lead Site Reliability Engineer (SRE) will help maintain and improve the reliability, performance, and security of software products and platforms. This role combines operational excellence with engineering rigor, with a strong emphasis on automation and Infrastructure-as-Code. The Lead SRE will be a core part of the team responsible for ensuring uptime and system health across services operating on both internal and client networks. The position will also contribute to technical development, platform enhancements, compliance initiatives, and the development of automated solutions for preventing, detecting, and remediating operational issues.

In This Role, You’ll:

  • Maintain the uptime, performance, and security of software products across internal and client environments.

  • Leverage automation and Infrastructure-as-Code to manage and scale infrastructure.

  • Respond to and resolve support requests during assigned SRE shifts.
  • Escalate support requests as necessary to meet established Service Level Agreements (SLAs).

  • Develop SRE Playbooks designed to automate the avoidance, detection, and remediation of operational issues.

  • Participate in code reviews with a focus on quality and security.
  • Contribute to accreditation and compliance initiatives.
  • Support platform capability enhancements and other technical development activities.

  • Help maintain system health and operational reliability across production services.
  • Support 24/7 production services through scheduled SRE duty rotations and shared after-hours on-call coverage.

Requirements

  • Demonstrated foundation in automation and Infrastructure-as-Code.
  • Demonstrated foundation in DevOps practices.
  • Demonstrated understanding of distributed systems.
  • Demonstrated experience working within cloud environments.
  • Ability to maintain and improve the reliability, performance, security, uptime, and overall health of software products and platforms.

  • Ability to manage and scale infrastructure through automation and Infrastructure as-Code.

  • Ability to respond to and resolve operational support requests and escalate issues as necessary to meet SLAs.

  • Experience working with technologies and environments that include: Kubernetes, Terraform, AWS, Go and/or Python

  • Active U.S. Government security clearance.
  • Ability to comply with applicable security obligations and procedural requirements.

Desired Skills:

  • Experience developing SRE playbooks for the automated avoidance, detection, or remediation of operational issues.

  • Experience performing code reviews for quality and security purposes.
  • Experience supporting accreditation and compliance initiatives.
  • Experience contributing to platform capability enhancements and other technical development activities.

  • Experience supporting 24/7 production services.
  • Experience supporting software products and platforms across both internal and client environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

3:27 min

Defining DevOps through its historical origins and foundational texts

Sonal Patil · LIVE

2:06 min

Elevating the QA engineering role for complex challenges

Ondřej Gróf Ondřej Gróf · World Congress 2026 Europe

Videos

See all

Related articles

See all