> Markdown version of [/jobs/ext/233142-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/233142-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Ddc It Services, LLC - **Location:** Scottsdale, AZ, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Agile Methodology, Amazon Web Services, Application Release Automation, Systems Engineering, ArcGIS (Software), Cloud Computing, Databases, Continuous Integration, Distributed Systems, Linux System Administration, Nagios, Performance Tuning, Reliability Engineering, Runbook, Software Vulnerability Management, Esri GIS (Software), Data Logging, Scripting, Enterprise Software Applications, Cloud Platform System, Software Security, Reliability of Systems, SC Clearance, Kubernetes, Information Technology, Devsecops - **Published:** May 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=4279cc65bf235db1 ## About the Role Do you have experience in Tooling?, Do you have a Master's degree?, * Active Secret clearance required. * Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field; Master's degree preferred. * Minimum of 8 years of experience supporting enterprise systems, cloud platforms, site reliability engineering, production engineering, systems engineering, or related technical roles. * Experience supporting AWS environments, including monitoring, performance tuning, troubleshooting, incident response, and operational sustainment. * Experience with Linux administration, scripting, and troubleshooting distributed applications in production environments. * Experience with containerized systems and orchestration platforms such as Kubernetes. * Experience supporting CI/CD pipelines, release automation, infrastructure-as-code, and operational reliability in Agile or DevSecOps environments. * Experience with monitoring, logging, and alerting tools used to support enterprise application performance and infrastructure visibility. * Strong analytical, troubleshooting, documentation, and communication skills, with the ability to translate operational issues into engineering improvements. * Ability to work effectively across cross-functional teams in a mission-focused DoD environment. Preferred * Experience supporting AWS Cloud One or other secure federal cloud environments. * Experience supporting geospatial or Esri-based platforms, including ArcGIS Enterprise or related technologies. * Familiarity with service reliability practices such as SLIs, SLOs, error budgets, incident postmortems, and capacity planning. * Experience with Risk Management Framework (RMF), STIG compliance, vulnerability remediation, and secure system hardening practices. * AWS, Kubernetes, or other relevant cloud or reliability engineering certifications. * Experience supporting technical refresh, platform modernization, or high-availability design initiatives in enterprise environments. ## Description The Site Reliability Engineer (SRE) / Subject Matter Expert (SME) - Computer Systems Engineer/Architect will provide senior-level reach-back expertise to support the reliability, scalability, performance, and operational resilience of the GEOMAP platform in secure cloud environments. This role focuses on improving service availability, monitoring, incident response, automation, and production stability across cloud-hosted and containerized systems supporting mission-critical geospatial capabilities for the U.S. Air Force. The Site Reliability Engineer will collaborate across development, DevSecOps, cloud, database, testing, and support teams to identify systemic issues, reduce operational risk, and implement engineering solutions that improve long-term platform reliability. *This position is contingent upon contract award.* Responsibilities * Provide senior-level engineering support to improve reliability, availability, performance, and maintainability of GEOMAP cloud-hosted systems and services. * Analyze production issues, recurring incidents, and operational trends to identify root causes and recommend durable corrective actions. * Support the design and implementation of monitoring, alerting, logging, and observability solutions across applications, infrastructure, and containerized services. * Develop and recommend automation approaches that reduce manual effort, improve deployment consistency, and increase system resilience. * Partner with software engineers, DevSecOps engineers, Kubernetes engineers, database engineers, and production support personnel to improve service health and release readiness. * Support incident response, problem management, service restoration, and post-incident reviews for high-priority operational issues. * Evaluate system performance, capacity, and scalability needs and provide recommendations for optimization and operational risk reduction. * Assist in defining service reliability objectives, operational metrics, and support models for sustained mission operations. * Contribute to infrastructure and platform engineering efforts involving cloud environments, CI/CD pipelines, container orchestration, and secure deployment patterns. * Support architecture reviews, technical assessments, and engineering analyses related to reliability, recoverability, and production operations. * Develop or refine runbooks, standard operating procedures, reliability engineering practices, and technical documentation. * Provide reach-back support for surge requirements, complex production investigations, and priority modernization or stabilization efforts as directed. * Performs other related duties as assigned. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [DevSecOps: Injecting Security into Mobile CI/CD Pipelines](https://www.wearedevelopers.com/videos/273-devsecops-injecting-security-into-mobile-ci-cd-pipelines) - [DevSecOps culture](https://www.wearedevelopers.com/videos/783-devsecops-culture) - [3 Key Steps for Optimizing DevOps Workflows](https://www.wearedevelopers.com/videos/962-3-key-steps-for-optimizing-devops-workflows) - [DevSecOps: Security in DevOps](https://www.wearedevelopers.com/videos/36-devsecops-security-in-devops) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering)