> Markdown version of [/jobs/ext/2045176-site-reliability-engineer-ii-govcloud](https://www.wearedevelopers.com/jobs/ext/2045176-site-reliability-engineer-ii-govcloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer II, GovCloud - **Company:** Medallia - **Location:** McLean, VA, United States - **Experience:** Experienced - **Salary:** $103,000.0 - $155,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Backup Devices, Cloud Computing, Computer Programming, Continuous Integration, Customer Data Management, Relational Databases, Linux, DevOps, Domain Name System (DNS), Federal Information Processing Standards (FIPS), Identity and Access Management, Subnetting, Python (Programming Language), Key Management, PostgreSQL, Routing, Octopus Deploy, Performance Tuning, Redis, Release Management, Reliability Engineering, Runbook, Software Engineering, Software Vulnerability Management, Data Logging, Google Cloud, Load Balancing, Cloud Platform System, Amazon Virtual Private Cloud (VPC), Git, Git Flow, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Github Enterprise, Apache Kafka, Terraform, Jenkins, Microservices - **Published:** August 13, 2026 - **Apply:** https://www.careerjet.com/job/usbff417e0fa61a66287092a513508b37e/eaa ## About the Role * Bachelor's degree or equivalent experience in Computer Science or a related field. * 2+ years of experience in Site Reliability Engineering, platform engineering, DevOps, or related production infrastructure roles (or equivalent hands-on production experience). * Must reside in the United States and be legally authorized to work in the US without sponsorship. * Production experience with: * Core services (IAM, compute, object storage, encryption/key management) and cloud networking on AWS (strongly preferred), Google Cloud (CGP), Azure, or a similar public cloud platform * Terraform or comparable infrastructure-as-code tools * Git and CI/CD pipelines * Linux and foundational systems concepts (networking, DNS, TLS/certificates) * PostgreSQL or another relational database - basic operations (queries, backups) * Familiarity with Kubernetes concepts, container orchestration, and microservices management * Programming and Automation: Proficiency in Python and/or Go experience to build automation scripts, operational tooling, and infrastructure services. * Incident & Change Management: Experience troubleshooting production incidents, conducting root-cause analysis(RCA), and following change management processes. * Experience participating in a production on-call rotation. * Experience troubleshooting complex technical issues and writing clear documentation, runbooks and incident post-mortems. Preferred Qualifications * Experience operating in FedRAMP, AWS GovCloud, or other regulated or compliance-heavy cloud environments. * Familiarity with security and compliance practices such as FIPS and vulnerability management. * Experience with observability and logging platforms in enterprise production environments. * Deep operational expertise with PostgreSQL (HA/replication, tuning, backup and recovery); familiarity with Redis and Kafka. * Experience supporting federal agencies or public-sector customers. * Experience with tools such as Jenkins, Argo CD, and GitHub Enterprise. * Strong collaboration skills and willingness to learn in a compliance-driven environment. ## Description We are growing our GovCloud team and looking for a Site Reliability Engineer II to help operate and improve Medallia's US public-sector cloud platform. You will support federal agencies and other regulated customers in a highly available, secure, and compliant environment built on AWS GovCloud and Kubernetes. This is a hybrid role based near Tysons, Virginia, with regular in-office collaboration and remote flexibility. You will be a hands-on engineer who helps operate and improve production systems, works closely with engineering and security teams, and grows toward owning larger systems while helping keep the platform reliable as we scale. Responsibilities * Build, operate, and improve highly available, secure cloud infrastructure on AWS, including networking, access management, Kubernetes clusters, DNS, certificates, and shared platform services. * Implement, operate and optimize AWS cloud networking, specifically managing VPCs, subnets and routing, security groups/NACLs, VPC endpoints/PrivateLink and load balancing. * Monitor, maintain and support production PostgreSQL - replication, backups and recovery, routine performance tuning, and upgrades - as part of the platform's data tier. * Monitor production systems, respond to incidents, and drive fixes that improve reliability and reduce repeat issues. * Develop and maintain Infrastructure-as-Code (primarily Terraform) and Kubernetes deployment workflows using Git, CI/CD, and GitOps practices. * Improve observability across metrics, logs, and uptime monitoring; help tune alerts and operational runbooks. * Partner with software engineering, security, and release management teams to deploy changes safely and resolve production issues. * Contribute to platform upgrades, security patching, and compliance-driven maintenance in a regulated cloud environment. * Participate in an on-call rotation for production support. * Use AI-assisted tooling responsibly, with attention to security, privacy, and customer data boundaries. * Learn the platform and grow your scope with mentorship from senior engineers. Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)