> Markdown version of [/jobs/ext/1448755-devops-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/1448755-devops-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOps Site Reliability Engineer (SRE) - **Company:** IT Veterans - **Location:** Washington, DC, United States - **Experience:** Expert - **Salary:** $150,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Performance Management, Microsoft Azure, Bash Shell, Cloud Computing, Continuous Availability, DevOps, Monitoring of Systems, Python (Programming Language), Reliability Engineering, Cloud Services, Software Deployment, Data Logging, Google Cloud, Software Troubleshooting, Multi-Cloud, Reliability of Systems, Containerization, Kubernetes - **Published:** July 26, 2026 - **Apply:** https://www.careerjet.com/job/us2095c7de79168e46964c78f377bee37b/eaa ## About the Role * TS/SCI security clearance. * Strong understanding of Site Reliability Engineering (SRE) principles and best practices. * Hands-on experience deploying and managing containerized applications using Kubernetes. * Experience administering and troubleshooting multi-cloud environments, including Google Cloud Platform (GCP), Microsoft Azure, and Amazon Web Services (AWS). * Experience implementing and maintaining enterprise monitoring, logging, and automated alerting solutions. * Proficiency with scripting and automation using languages such as Python, Bash, or similar technologies., * Passion for building and maintaining highly reliable, mission-critical systems with demanding uptime requirements. * Ability to remain composed and methodical while responding to high-priority production incidents. * Strong troubleshooting, root cause analysis, and diagnostic skills with a focus on rapid issue resolution. * A continuous improvement mindset with an emphasis on automation and eliminating repetitive operational tasks. * Experience supporting secure, cloud-native environments within government or defense organizations is a plus. ## Description IT Veterans is seeking a DevOps Site Reliability Engineer (SRE) to support the reliability, performance, and operational stability of a mission-critical enterprise platform. This position plays a vital role in ensuring continuous availability across multi-cloud environments while supporting software deployments, infrastructure monitoring, incident response, and system automation. The ideal candidate will help bridge development and operations by implementing reliable deployment practices, building robust monitoring capabilities, and rapidly responding to production issues to maintain the required 99.9% platform availability for critical Department of Defense (DoD) systems., * Monitor the health and performance of enterprise infrastructure through continuous system monitoring and automated telemetry to support the required 99.9% platform uptime. * Participate in an on-call rotation and respond to major incidents or platform outages within one hour of notification, executing rapid troubleshooting and system stabilization activities. * Develop, maintain, and enhance automation scripts and internal tools that streamline diagnostics, health checks, and routine operational tasks. * Design and maintain dashboards that provide real-time visibility into platform health, including uptime, API performance, incident status, and other key operational metrics. * Coordinate directly with Cloud Service Providers (CSPs) during infrastructure outages or service disruptions to expedite issue resolution. * Continuously assess system reliability, logging, monitoring, and overall architecture, providing recommendations that improve scalability, resiliency, and operational efficiency. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Old tools, new tricks](https://www.wearedevelopers.com/videos/1916-old-tools-new-tricks) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)