Site Reliability Engineer (SRE) - Day Shift

Peraton Inc
United States
3 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$104,000.0 - $166,000.0
Working hours
Shift work

Tech stack

Amazon Web Services Microsoft Azure Backup Devices Bash Shell Cloud Engineering Continuous Integration Noise Reduction Linux DevOps Disaster Recovery Failover Python (Programming Language)
+21 more
Windows Servers OpenShift Windows PowerShell Red Hat Enterprise Linux Reliability Engineering Ansible Datadog Scripting Performance Testing Delivery Pipeline Grafana Technical Debt Gitlab Gitlab-ci Kubernetes Deployment Automation Terraform Splunk Ansible Tower Dynatrace Jenkins

Job description

Peraton is seeking a Site Reliability Engineer (SRE) to join a team respobsible for the operational reliability of production systems running in AWS Commercial and AWS GovCloud environments. The ideal candidate has working knowledge of deploying and managing system components in Azure and GCP. The primary production workload runs on Red Hat OpenShift Service on AWS (ROSA)

The SRE partners closely with platform engineers, the security team, and application developers to ensure the infrastructure services are reliable, available, and deployed in a way that meets both developer and security requirements., * Operate and maintain production infrastructure services and applications, ensuring availability and reliability, performance, security, and operational health.

  • Monitor services and applications using defined SLIs, SLOs, dashboards, alerts, and other observability tools; continuously improve the detection, diagnosis, and resolution of operational issues.
  • Partner with application teams to define application observability requirements and implement appropriate metrics, logs, traces, dashboards, and alerts into the organization’s observability tooling.
  • Manage production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions.
  • Execute application and infrastructure releases through established deployment pipelines, including promotion through staging and production, validation, rollback, and release-related troubleshooting.
  • Manage the operational lifecycle of deployed infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes.
  • Assess and improve service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing.
  • Identify and address reliability risks and operational technical debt by using reliability metrics, incident trends, capacity data, and service health indicators to prioritize improvements.
  • Automate operational activities using an everything-as-code approach to improve consistency, repeatability, testing, deployment, recovery, and operational efficiency.
  • Collaborate with platform engineering and application teams to identify operational requirements, provide feedback on reusable infrastructure building blocks, and continuously improve the reliability and operability of the environment.

Requirements

  • Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance.
  • Bachelor’s Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience.
  • 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering.
  • Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms
  • Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower.
  • Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation.
  • Proficient in Linux and Windows Server administration
  • Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry.
  • Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction.
  • Scripting/automation proficiency in Python, Bash, PowerShell, or Go.
  • Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).

Preferred Qualifications:

  • AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification
  • Red Hat Certified Specialist in ROSA, Red Hat Certified System Administrator in OpenShift
  • Azure Administrator Associate, GCP Associate Cloud Engineer certification
  • Dynatrace Associate, Datadog Log Management Fundamentals certification
  • GitLab CI/CD Associate certification, Certified Jenkins Engineer (CJE)
  • Terraform Associate certification

Benefits & conditions

Target Salary Range: $104,000 - $166,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.

Benefits Statement: Peraton offers eligible employees a variety of benefits including medical, dental, vision, life, health savings account, short/long term disability, EAP, parental leave, 401(k), paid time off (PTO) for vacation, and company paid holidays. A full listing of available benefits can be viewed at https://www.careers.peraton.com/benefits.

About the company

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees solve the most daunting challenges that our customers face. Visit peraton.com to learn how we’re keeping people around the world safe and secure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all