> Markdown version of [/jobs/ext/2059150-staff-site-reliability-engineer-govcloud](https://www.wearedevelopers.com/jobs/ext/2059150-staff-site-reliability-engineer-govcloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer, GovCloud - **Company:** Medallia - **Location:** McLean, VA, United States - **Experience:** Expert - **Salary:** $158,500.0 - $230,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Computing Platforms, Backup Devices, Computer Programming, Continuous Integration, Customer Data Management, Relational Databases, Linux, DevOps, Domain Name System (DNS), Federal Information Processing Standards (FIPS), Monitoring of Systems, Identity and Access Management, Subnetting, Python (Programming Language), Key Management, PostgreSQL, Routing, Octopus Deploy, Performance Tuning, Redis, Release Management, Reliability Engineering, Runbook, Software Engineering, Management of Software Versions, Software Vulnerability Management, Data Logging, Load Balancing, System Availability, Amazon Virtual Private Cloud (VPC), Git, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Github Enterprise, Apache Kafka, Terraform, Jenkins - **Published:** August 14, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88030184/1 ## About the Role * Must reside in the United States and be legally authorized to work in the US without sponsorship. * Bachelor's degree or equivalent experience in Computer Science or a related field. * 8+ years of experience in Site Reliability Engineering, platform engineering, DevOps, or related production infrastructure roles - or 5+ years with demonstrated Staff-level scope (technical leadership, cross-team delivery, incident ownership, and platform/IaC ownership). * Production experience with: + Kubernetes + AWS core services (IAM, compute, object storage, encryption/key management) and AWS cloud networking (VPC design, routing, security groups, load balancing, VPC endpoints/PrivateLink, Transit Gateway, DNS) + Terraform or comparable infrastructure-as-code tools + Git and CI/CD pipelines + Linux and foundational systems concepts (networking, DNS, TLS/certificates) + PostgreSQL (or comparable relational databases) in production - replication, backups, and performance tuning * Programming and Automation: Proficiency in Python and/or Go experience to build automation scripts, operational tooling, and infrastructure services. * Incident & Change Management: Experience troubleshooting production incidents, conducting root-cause analysis(RCA), and following change management processes. * Experience participating in a production on-call rotation. * Experience troubleshooting complex technical issues and writing clear documentation, runbooks and incident post-mortems. Preferred Qualifications * Experience operating in FedRAMP, AWS GovCloud, or other regulated or compliance-heavy cloud environments. * Familiarity with security and compliance practices such as FIPS, vulnerability management, and controlled production change processes. * Experience with observability and logging platforms in enterprise production environments. * Deep operational expertise with PostgreSQL (HA/replication, tuning, backup and recovery); familiarity with data/platform technologies such as Redis and Kafka. * Experience supporting federal agencies or public-sector customers. * Exposure to government networking and security requirements. * Experience with tools such as Jenkins, Argo CD, and GitHub Enterprise. * Excellent collaboration skills and a strong willingness to learn. ## Description * Design, build, and operate highly available, secure cloud infrastructure on AWS, including networking, identity/access management, Kubernetes clusters, DNS, certificates, and shared platform services. * Design and operate AWS cloud networking end-to-end - VPC architecture, subnetting and routing, security groups/NACLs, VPC endpoints/PrivateLink, Transit Gateway, load balancing, and DNS - for secure, segmented, highly available connectivity. * Operate and tune production PostgreSQL - high availability and replication, backups and recovery, query and performance optimization, version upgrades, and capacity planning - as part of the platform's data tier. * Ensure the reliability and availability of Medallia applications and infrastructure by monitoring systems, responding to incidents, and eliminating recurring operational problems. * Develop and maintain Infrastructure-as-Code (primarily Terraform) and Kubernetes deployment workflows using Git, CI/CD, and modern GitOps practices. * Improve observability across metrics, logs, and uptime monitoring; tune alerting to reduce noise and speed up diagnosis. * Partner with software engineering, security, release management, and customer-facing teams to deploy changes safely, resolve production issues, and improve operability. * Lead or contribute to platform upgrades, security patching, and compliance-driven maintenance in a regulated cloud environment. * Participate in an on-call rotation and help improve incident response, communication, and post-incident follow-through. * Document systems and operational procedures clearly so others can run and improve the platform. * Use AI-assisted tooling responsibly, with attention to security, privacy, and customer data boundaries. * Mentor engineers and help raise engineering standards as the GovCloud platform and team grow. Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)