> Markdown version of [/jobs/ext/2843529-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2843529-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** TEKSYSTEMS INC. - **Location:** Joint Base Pearl Harbor-Hickam, HI, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Active Directory, Active Directory Federation Services, Agile Methodology, Amazon Web Services, Component-Based Software Engineering, Confluence, JIRA, Microsoft Azure, Bash Shell, Cloud Computing, Computer Networks, Continuous Integration, Dynamic Host Configuration Protocol, DevOps, Domain Name System (DNS), Python (Programming Language), Load Testing, Microsoft SQL Server, Windows Servers, OpenShift, Public Key Infrastructure, Windows PowerShell, Scrum Methodology, Reliability Engineering, Ansible, Software Deployment, Software Engineering, Strategies of Testing, Data Logging, Scripting, Cloud Platform System, Performance Testing, Reliability of Systems, Gitlab, Cloudformation, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Performance Monitor, Bitbucket, Puppet, Software Coding, Terraform, Splunk, Docker, Jenkins - **Published:** September 11, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9152544/site-reliability-engineer ## About the Role BS + 2-4 years of experience (or MS + <2 years). Active DoD Secret clearance and DoD 8570 IAT II certification. Ability to support classified environments; 100% onsite with SIPRNet access. Experience with Active Directory, GPOs, DNS, DHCP, and Windows Server (2016-2025). Experience with SQL Server (2019-2025). Linux/Unix administration and scripting (PowerShell, Python, Bash). Experience with CI/CD tools (Jenkins, GitLab) and DevOps/SRE environments. Hands-on experience with Jira, Confluence, Bitbucket, Azure DevOps, and Ansible. Experience with OpenShift, Kubernetes, Docker, AWS, and Azure. Familiarity with Infrastructure as Code (Terraform, CloudFormation, Ansible, Chef, Puppet). Knowledge of RMF, DISA STIGs, and application administration/integration. Strong collaboration skills in Agile environments., NGEN-NMCI program experience. ITIL, Scrum Master, SAFe, or similar certifications. Experience with Terraform, Ansible, or CloudFormation. Familiarity with MIM/FIM, Delinea, ADFS, and PKI. ## Description The SRE will also develop and execute tests focused on system resilience, performance under load, and failure scenarios. They will work in tandem with other Site Reliability Engineers (SREs) and development teams to create automated testing frameworks that simulate real-world conditions that validate system behavior under normal and stress conditions, ensuring our services are resilient and meet established service level objectives (SLOs). Your work will contribute to the development of robust and scalable services that operate reliably in production. Your responsibilities will include maintaining complex computer systems by writing code to automate software releases, monitor systems, and detect and fix problems before users even know there is an issue. You will use these skills to improve site performance and overall reliability. The SRE-IDAM role is responsible for supporting, migrating, automation and optimization of software development and deployment process, infrastructure as code, and contribute to the overall maturity of the Site Reliability Engineering program as well as supporting the creation, maintenance, update, modernization, and refresh the capabilities and components of the Navy's Enterprise Network. The SRE-IDAM resource provides technical leadership and knowledge of related tasks and coordinates with project managers, customers, stakeholders, and engineers to support ongoing activities as well as new projects to maintain, transform, and modernize the Navy Enterprise Network. * Test, maintain (patching, STIGing, and upgrading), troubleshoot, develop, and deliver solutions associated with Active Directory, Azure, Delinea, Ansible, Microsoft Identity Manager (MIM), Active Directory Federation Services (ADFS), DHCP, DNS, WINS, GPOs & PKI. * Work alongside the development and operations teams to ensure speedy and reliable software deployments, monitor systems, and improve overall reliability of the platform. In addition, as you discover and document system bugs, you have the motivation to go off and fix them yourself. * Develop features utilize the AI coding tool and repository of scripts to automate, scale, test, and secure the cloud infrastructure and the pipelines. * Enhance performance monitoring of the various systems via Splunk or other dashboard reporting tools * Identify performance bottlenecks and optimize the performance of cloud infrastructure * Contribute to continuing our SRE journey by suggesting ways to improve engineering build, maintenance, automation and reliability across the platform with SRE/DevOps tools and Infrastructure-as-Code. * Develop and code high-quality pipeline automation workflows to support inside and outside the cloud platform that are appropriate for business and technology strategies. * Develop and execute test strategies that simulate real-world failure scenarios, including network disruptions, hardware failures, and system overloads. * Create, script, and run performance tests to measure system behavior under varying levels of load and traffic. Identify bottlenecks, performance degradation, and areas for optimization. * Design, implement, and maintain automated test suites for infrastructure and application components. Ensure that testing is integrated into the CI/CD pipeline to validate system reliability with every release. * Build automated systems for continuous performance testing, stress testing, and load testing. * Work closely with SREs, developers, and operations teams to define reliability goals and develop appropriate testing strategies to validate those goals. * Ensure that new services and features undergo thorough testing for performance, reliability, and failure recovery before deployment to production. * Validate that monitoring, logging, and alerting mechanisms are functioning correctly by testing systems under failure conditions. * Ensure that Service Level Indicators (SLIs) and Service Level Objectives (SLOs) are accurately measured and tracked through automated testing frameworks. * Resolve most conflicts between timeline, budget, and scope independently but intuitively raise sophisticated or consequential issues to senior management. * Must be willing to work nights, weekends, and provide on-call support as needed. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)