> Markdown version of [/jobs/ext/3052553-sre-government-cloud-operations](https://www.wearedevelopers.com/jobs/ext/3052553-sre-government-cloud-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE - Government Cloud Operations - **Company:** Cato Networks - **Location:** Dallas, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Bash Shell, Software as a Service, Cloud Computing, Configuration Management, Linux, Disaster Recovery, Distributed Systems, Python (Programming Language), Reliability Engineering, Cloud Services, Prometheus, Zero Trust Network Access, Software Engineering, Software Vulnerability Management, Data Logging, Grafana, Containerization, Kubernetes, Infrastructure Automation Frameworks, Terraform - **Published:** September 24, 2026 - **Apply:** https://www.juju.com/job/15_75bbf3ff ## About the Role * 7+ years of experience in Site Reliability Engineering, Production Engineering, Cloud Operations, or Infrastructure Engineering. * Hands-on experience operating cloud infrastructure in regulated environments such as FedRAMP Moderate/High, DoD IL4/IL5, or equivalent, including AWS GovCloud or other isolated government cloud environments. * Experience supporting cloud authorization efforts (ATO) and sustaining environments post-authorization through continuous monitoring, including FedRAMP monthly reporting, vulnerability tracking, and control assessment activities. * Strong knowledge of NIST 800-53 controls, vulnerability remediation SLAs, secure configuration management, and audit evidence generation. * Deep experience with Infrastructure as Code (Terraform preferred), GitOps workflows, and secure CI/CD pipelines, including container hardening and image security practices. * Proficiency in Python, Go, or Bash for operational automation and tooling development. * Proficiency with cloud-native technologies including Kubernetes, Prometheus, and Grafana, along with a solid understanding of Linux/Unix operating systems. * Experience supporting production operations for SaaS, cloud service provider, or multi-tenant platforms at scale. * Ability to communicate operational risk and compliance posture clearly to both technical and non-technical stakeholders. Preferred Qualifications * Experience working directly with 3PAOs, auditors, or compliance assessors during authorization and continuous monitoring cycles. * Familiarity with STIG implementation across Kubernetes, Linux systems, and container runtimes. * Understanding of Zero Trust architectures and secure access platforms. * Experience with operational resilience exercises and disaster recovery validation. ## Description We're seeking a Senior Site Reliability Engineer with hands-on experience building and sustaining regulated cloud platforms through FedRAMP High / IL4 operational lifecycles, including continuous monitoring and post-ATO operational management. In this critical role, you will support our growing operations, network, and systems environments. You will play a pivotal role in administering internal platforms while participating in key architectural and operational decisions. This position offers the opportunity to innovate, establish best-practice processes, and continuously improve the reliability, security, and compliance posture of our regulated cloud environments. Responsibilities * Own production operations for mission-critical services, including availability, latency, scalability, and operational health across complex distributed systems. * Design, build, and operate highly available cloud infrastructure supporting regulated environments, including FedRAMP High / IL4+ deployments. * Lead major incident response, root cause analysis, and postmortem remediation; drive operational maturity through change governance, disaster recovery testing, and service resiliency programs. * Operationalize compliance requirements, including NIST 800-53 controls and STIG baselines, across Kubernetes platforms, Linux systems, container runtimes, and cloud infrastructure. * Support regulated environment readiness through audit preparation, evidence collection, vulnerability management, configuration management, and continuous monitoring activities. * Develop automation and tooling to continuously assess and maintain platform compliance posture; contribute to immutable, reproducible infrastructure patterns that simplify regulatory sustainment. * Implement and maintain secure CI/CD pipelines and infrastructure-as-code practices aligned with security and compliance requirements. * Improve observability across infrastructure and applications through metrics, logging, tracing, and alerting; integrate compliance telemetry and configuration auditing into operational workflows. * Partner with Security, Compliance, and Engineering teams to improve service reliability, deployment safety, and operational maturity throughout the software lifecycle. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [The Geometry of Incidents: Connecting User Impact to Architecture](https://www.wearedevelopers.com/magazine/764-the-geometry-of-incidents-connecting-user-impact-to-architecture) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.](https://www.wearedevelopers.com/magazine/693-events-like-rsac-get-you-cisos-developers-decide-what-actually-gets-deployed) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)