Site Reliability Engineer

Eliassen Group
Norfolk, VA, United States
1 day ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$124,800.0 - $135,200.0
Working hours
Regular working hours

Tech stack

Active Directory Active Directory Federation Services Automation of Tests Microsoft Azure Cloud Computing Computer Networks Continuous Integration Dynamic Host Configuration Protocol DevOps Distributed Systems Domain Name System (DNS) Identity and Access Management
+14 more
Load Testing Public Key Infrastructure Reliability Engineering Site Reliability Engineering Practices Ansible Software Deployment Strategies of Testing Data Logging System Availability Delivery Pipeline SC Clearance Performance Monitor Windows Services Splunk

Job description

Our client seeks a Site Reliability Engineer to focus on reliability, performance, and scalability of distributed systems supporting the Navy Enterprise Network. The role will design and automate tests for resilience, performance under load, and failure scenarios. It will collaborate with SRE and development teams to integrate automated testing into CI/CD, validate SLOs, and improve monitoring and alerting. The position will also automate releases, monitor systems, and remediate issues proactively to improve site performance and reliability. This role supports identity and access management platforms and contributes to modernization of the enterprise network., * Test, maintain, troubleshoot, and upgrade solutions for Active Directory, Azure, Delinea, Ansible, Microsoft Identity Manager, ADFS, DHCP, DNS, WINS, GPOs, and PKI, including patching and STIG compliance.

  • Collaborate with development and operations to enable reliable deployments, monitor systems, and improve platform reliability. Identify, document, and resolve system bugs.
  • Develop features and scripts using AI coding tools to automate, scale, test, and secure cloud infrastructure and pipelines.
  • Enhance performance monitoring using Splunk or other dashboards.
  • Identify and remediate performance bottlenecks in cloud infrastructure.
  • Advance SRE practices for build, maintenance, automation, and reliability using DevOps tools and Infrastructure as Code.
  • Design and code pipeline automation workflows across cloud and on-prem environments aligned to business and technology strategies.
  • Develop and execute test strategies simulating failures such as network disruptions, hardware faults, and overloads.
  • Create, script, and run performance tests to assess system behavior under varying load and traffic. Identify degradation and optimization areas.
  • Design, implement, and maintain automated test suites for infrastructure and applications. Integrate testing into CI/CD.
  • Build automated systems for continuous performance, stress, and load testing.
  • Partner with SREs, developers, and operations to define reliability goals and validate them through testing.
  • Ensure new services and features undergo performance, reliability, and recovery testing before production deployment.
  • Validate monitoring, logging, and alerting by testing under failure conditions.
  • Ensure SLIs and SLOs are measured and tracked via automation.
  • Resolve most timeline, budget, and scope conflicts independently and escalate consequential issues appropriately.
  • Provide night, weekend, and on-call support as needed.

Requirements

Due to federal security clearance requirements, applicant must be a United States Citizen with an active Secret clearance. Due to client requirements, applicants must be willing and able to work on a w2 basis. For our w2 consultants, we offer a great benefits package that includes Medical, Dental, and Vision benefits, 401k with company matching, and life insurance., * Experience with identity platforms and Windows services including Active Directory, ADFS, GPOs, PKI, DHCP, DNS, and WINS.

  • Hands-on experience with Azure, Ansible, Delinea, and Microsoft Identity Manager.
  • Proficiency building CI/CD-integrated automated tests for performance, stress, and failure scenarios.
  • Experience with performance monitoring and observability using Splunk or similar tools.
  • Ability to identify and remediate infrastructure performance bottlenecks.
  • Experience developing Infrastructure as Code and pipeline automation.
  • Ability to collaborate across SRE, development, and operations teams to meet SLOs.
  • Willingness to support alternate schedules and on-call rotations.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all