> Markdown version of [/jobs/ext/3195621-cloud-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/3195621-cloud-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Site Reliability Engineer (SRE) - **Company:** Everforth Ecs - **Location:** Arlington, VA, United States - **Experience:** Expert - **Salary:** $130,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Cloud Computing, Continuous Integration, Software Debugging, Python (Programming Language), Uptime, Reliability Engineering, Prometheus, Software Engineering, Datadog, Data Logging, Cloud Platform System, Grafana, Build Management, Kubernetes, Information Technology, Terraform, Splunk, Jenkins - **Published:** September 16, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=dfbfa296f04bb08e ## About the Role * Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent practical experience) * 5+ years of SRE experience (or equivalent), with demonstrated technical leadership * 10 years of general work experience * Track record building self-healing/auto-remediating systems, not just dashboards * Jenkins experience * Expert AWS knowledge, GovCloud experience strongly preferred * Deep Kubernetes and Terraform expertise at production scale * Strong software engineering background (Python and/or Go) * Experience operating observability platforms (Grafana, Splunk, Prometheus, Loki, etc.) * Proven incident command and postmortem experience * Strong communication skills across technical and federal leadership audiences * Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable) * US Citizenship ## Description We believe the job of an SRE is to engineer the cloud to run itself. That means writing software and automation that lets systems detect and recover from failure on their own, rather than relying on someone to notice an alert and manually fix it. When something breaks, self-healing comes first, deep root-cause debugging happens after service is restored, not instead of it. We're looking for someone who automates the operational task by default, not documents the runbook for doing it by hand., This role owns reliability and operational readiness for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You'll define what "reliable enough" looks like for our services, build the automation that gets us there, and do it all on an infrastructure-as-code (IaC) foundation., Self-Healing Operations * Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention * Shift the team's posture from "is it running, how do we fix it" to "how do we make it fix itself" * Automate service restoration first; investigate root cause after Uptime Goals & Reliability * Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them * Use live metrics to decide what's "reliable enough" and where to invest next Infrastructure * Enforce infrastructure-as-code and configuration-as-code, no manual tech change * Own Terraform standards and reusable modules adopted across programs * Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling) * Set CI/CD and pipeline-as-code standards, including progressive delivery Observability & Incidents * Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct * Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation Collaboration & Leadership * Partner with development and contractor teams leads to embed reliability and automation across the software * Mentor engineers toward this same automation-first philosophy * Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation Salary Range: $130,000 - $180,000 ## Related Videos - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [GitLab CI pipelines for a whole company](https://www.wearedevelopers.com/videos/143-gitlab-ci-pipelines-for-a-whole-company) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)