> Markdown version of [/jobs/ext/3584411-site-reliability-engineer-ii](https://www.wearedevelopers.com/jobs/ext/3584411-site-reliability-engineer-ii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer II - **Company:** Medallia - **Location:** McLean, VA, United States - **Experience:** Experienced - **Salary:** $103,500.0 - $150,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Systems Engineering, Build Automation, Bash Shell, Cloud Computing, Continuous Integration, DevOps, Distributed Systems, Domain Name System (DNS), Python (Programming Language), Linux System Administration, Networking Basics, Routing, Reliability Engineering, Software Tools, Prometheus, Software Engineering, Scripting, Transport Layer Security, Load Balancing, Grafana, Reliability of Systems, HybridCloud, Git Flow, Kubernetes, Infrastructure Automation Frameworks, Terraform, Oracle Cloud Infrastructure, Service Stack, Programming Languages - **Published:** October 4, 2026 - **Apply:** https://www.juju.com/job/21_01a0ef03-aadd-7013-9033-0ceb24676e72 ## About the Role * 2+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Operations, or related roles. * Demonstrated experience supporting production environments running on Kubernetes or other containerized platforms. * Demonstrated experience with cloud infrastructure platforms such as AWS, OCI, or GCP. * Demonstrated experience with Linux systems administration and troubleshooting. * Demonstrated experience with scripting or programming languages such as Python, Bash, or Go. * Familiarity with CI/CD pipelines and Git-based workflows. * Demonstrated understanding of networking fundamentals including DNS, load balancing, TLS/SSL, and routing concepts. * Demonstrated experience troubleshooting distributed systems and production incidents. * Ability to participate in an on-call rotation supporting production systems. * Fluency in English, both oral and written. Preferred Qualifications * Experience with GitOps and tools such as ArgoCD. * Experience with infrastructure-as-code tools such as Terraform. * Familiarity with observability platforms such as Prometheus, Grafana, Loki, or OpenTelemetry. * Experience operating services in hybrid-cloud or multi-region environments. * Understanding of release strategies such as rolling deployments, canary releases, or blue/green deployments. * Familiarity with incident management and operational best practices. * Exposure to security and compliance concepts in production environments. * Experience using AI-assisted development, automation, or operational tooling to improve engineering productivity and service reliability. * Demonstrated passion for automation, process improvement, and operational efficiency. * Strong communication and collaboration skills. ## Description As an SRE II, you will help operate and improve the reliability, scalability, and performance of services running across Kubernetes-based environments in cloud and hybrid infrastructure. You will work closely with software engineering teams to build automation, improve operational excellence, and support production services used globally by Medallia customers. We are looking for engineers who enjoy solving complex technical problems, automating repetitive tasks, improving system reliability, and learning modern cloud-native technologies in a fast-paced environment. We value engineers who actively seek opportunities to improve scalability and operational efficiency through automation, AI-assisted engineering workflows, and continuous process improvement. Please note this role participates in a rotating on-call schedule supporting production systems and services. Engineering Leverage At Medallia, we hire engineers who scale systems, teams, and outcomes through automation, platform thinking, and AI-assisted engineering. We value engineers who challenge manual processes, reduce operational toil, and create reusable solutions that improve reliability and productivity for the broader engineering organization. Successful engineers do not simply solve problems-they eliminate recurring problems through automation, simplification, and self-service capabilities. Responsibilities * Collaborate with software engineering teams to improve application reliability, scalability, and operational maturity. * Operate and support production services running in Kubernetes environments. * Troubleshoot and resolve infrastructure and application issues across the full technology stack. * Build automation and tooling to reduce operational overhead and eliminate manual work. * Leverage AI-assisted engineering tools and automation platforms to accelerate troubleshooting, improve productivity, and reduce operational toil. * Identify opportunities to streamline operational processes through automation, AI-enabled workflows, and self-service solutions. * Create reusable solutions, tooling, and operational improvements that increase engineering leverage across the team. * Support CI/CD and GitOps-based deployment workflows. * Develop and maintain infrastructure-as-code configurations and operational tooling. * Monitor system health, availability, and performance using observability and alerting platforms. * Participate in incident response, root cause analysis, and operational improvements. * Continuously improve reliability, deployment processes, and operational standards. Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite. ## Related Videos - [Shifting Stress to Progress— Understanding DevOps to do DevOps Better](https://www.wearedevelopers.com/videos/268-shifting-stress-to-progress-understanding-devops-to-do-devops-better) - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)