Site Reliability Engineer

TEKSYSTEMS INC.
Joint Base Pearl Harbor-Hickam, HI, United States
10 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Shift work

Tech stack

Active Directory Active Directory Federation Services Agile Methodology Amazon Web Services Component-Based Software Engineering Confluence JIRA Microsoft Azure Bash Shell Cloud Computing Computer Networks Continuous Integration
+34 more
Dynamic Host Configuration Protocol DevOps Domain Name System (DNS) Python (Programming Language) Load Testing Microsoft SQL Server Windows Servers OpenShift Public Key Infrastructure Windows PowerShell Scrum Methodology Reliability Engineering Ansible Software Deployment Software Engineering Strategies of Testing Data Logging Scripting Cloud Platform System Performance Testing Reliability of Systems Gitlab Cloudformation Kubernetes Infrastructure Automation Frameworks Information Technology Performance Monitor Bitbucket Puppet Software Coding Terraform Splunk Docker Jenkins

Job description

The SRE will also develop and execute tests focused on system resilience, performance under load, and failure scenarios. They will work in tandem with other Site Reliability Engineers (SREs) and development teams to create automated testing frameworks that simulate real-world conditions that validate system behavior under normal and stress conditions, ensuring our services are resilient and meet established service level objectives (SLOs). Your work will contribute to the development of robust and scalable services that operate reliably in production. Your responsibilities will include maintaining complex computer systems by writing code to automate software releases, monitor systems, and detect and fix problems before users even know there is an issue. You will use these skills to improve site performance and overall reliability.

The SRE-IDAM role is responsible for supporting, migrating, automation and optimization of software development and deployment process, infrastructure as code, and contribute to the overall maturity of the Site Reliability Engineering program as well as supporting the creation, maintenance, update, modernization, and refresh the capabilities and components of the Navy’s Enterprise Network. The SRE-IDAM resource provides technical leadership and knowledge of related tasks and coordinates with project managers, customers, stakeholders, and engineers to support ongoing activities as well as new projects to maintain, transform, and modernize the Navy Enterprise Network.

  • Test, maintain (patching, STIGing, and upgrading), troubleshoot, develop, and deliver solutions associated with Active Directory, Azure, Delinea, Ansible, Microsoft Identity Manager (MIM), Active Directory Federation Services (ADFS), DHCP, DNS, WINS, GPOs & PKI.
  • Work alongside the development and operations teams to ensure speedy and reliable software deployments, monitor systems, and improve overall reliability of the platform. In addition, as you discover and document system bugs, you have the motivation to go off and fix them yourself.
  • Develop features utilize the AI coding tool and repository of scripts to automate, scale, test, and secure the cloud infrastructure and the pipelines.
  • Enhance performance monitoring of the various systems via Splunk or other dashboard reporting tools
  • Identify performance bottlenecks and optimize the performance of cloud infrastructure
  • Contribute to continuing our SRE journey by suggesting ways to improve engineering build, maintenance, automation and reliability across the platform with SRE/DevOps tools and Infrastructure-as-Code.
  • Develop and code high-quality pipeline automation workflows to support inside and outside the cloud platform that are appropriate for business and technology strategies.
  • Develop and execute test strategies that simulate real-world failure scenarios, including network disruptions, hardware failures, and system overloads.
  • Create, script, and run performance tests to measure system behavior under varying levels of load and traffic. Identify bottlenecks, performance degradation, and areas for optimization.
  • Design, implement, and maintain automated test suites for infrastructure and application components. Ensure that testing is integrated into the CI/CD pipeline to validate system reliability with every release.
  • Build automated systems for continuous performance testing, stress testing, and load testing.
  • Work closely with SREs, developers, and operations teams to define reliability goals and develop appropriate testing strategies to validate those goals.
  • Ensure that new services and features undergo thorough testing for performance, reliability, and failure recovery before deployment to production.
  • Validate that monitoring, logging, and alerting mechanisms are functioning correctly by testing systems under failure conditions.
  • Ensure that Service Level Indicators (SLIs) and Service Level Objectives (SLOs) are accurately measured and tracked through automated testing frameworks.
  • Resolve most conflicts between timeline, budget, and scope independently but intuitively raise sophisticated or consequential issues to senior management.
  • Must be willing to work nights, weekends, and provide on-call support as needed.

Requirements

BS + 2-4 years of experience (or MS + <2 years). Active DoD Secret clearance and DoD 8570 IAT II certification. Ability to support classified environments; 100% onsite with SIPRNet access. Experience with Active Directory, GPOs, DNS, DHCP, and Windows Server (2016-2025). Experience with SQL Server (2019-2025). Linux/Unix administration and scripting (PowerShell, Python, Bash). Experience with CI/CD tools (Jenkins, GitLab) and DevOps/SRE environments. Hands-on experience with Jira, Confluence, Bitbucket, Azure DevOps, and Ansible. Experience with OpenShift, Kubernetes, Docker, AWS, and Azure. Familiarity with Infrastructure as Code (Terraform, CloudFormation, Ansible, Chef, Puppet). Knowledge of RMF, DISA STIGs, and application administration/integration. Strong collaboration skills in Agile environments., NGEN-NMCI program experience. ITIL, Scrum Master, SAFe, or similar certifications. Experience with Terraform, Ansible, or CloudFormation. Familiarity with MIM/FIM, Delinea, ADFS, and PKI.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all