Senior Site Reliability Engineer

Medallia
McLean, VA, United States
22 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$128,500.0 - $190,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Bash Shell Software as a Service Cloud Computing Continuous Integration Distributed Systems Domain Name System (DNS) Python (Programming Language) Linux System Administration Networking Basics Routing
+15 more
Reliability Engineering Prometheus Software Engineering Data Logging Transport Layer Security Load Balancing Delivery Pipeline Grafana HybridCloud Containerization Kubernetes Deployment Automation Terraform Oracle Cloud Infrastructure Golang

Job description

As a Senior Site Reliability Engineer, you will play a key role in designing, operating, and evolving the platforms and services that power Medallia’s global production environment. You will work across engineering teams to improve reliability, scalability, performance, and operational maturity while driving automation and platform improvements at scale. This role is expected to provide technical leadership, influence engineering best practices, and help shape the future direction of our cloud-native infrastructure and operational strategy. We are looking for engineers who think beyond day-to-day operations and continuously seek ways to increase engineering leverage. Successful candidates will act as force multipliers by building automation, self-service capabilities, platform solutions, and AI-assisted workflows that enable teams to operate more efficiently and reliably at scale. Please note this role participates in a rotating on-call schedule supporting production systems and services. Engineering Leverage At Medallia, we believe great engineers amplify the impact of themselves and those around them. Successful Senior SREs build systems, platforms, standards, and automation that enable multiple teams to move faster, operate more reliably, and scale efficiently. As AI capabilities continue to evolve, Senior SREs are expected to evaluate, adopt, and promote AI-assisted engineering practices that improve productivity, accelerate delivery, and reduce operational burden across the organization. Responsibilities:

  • Design, build, and operate highly available, scalable, and secure production platforms.
  • Partner with software engineering teams to improve application reliability, scalability, performance, and operational readiness.
  • Lead complex incident investigations, root cause analyses, and reliability improvement initiatives.
  • Design and implement automation, self-service capabilities, and platform solutions that reduce operational toil.
  • Leverage AI-assisted engineering tools and automation platforms to accelerate troubleshooting, improve productivity, and reduce operational overhead.
  • Identify opportunities to streamline operational processes through automation, AI-enabled workflows, and platform engineering practices.
  • Drive adoption of SRE principles, reliability standards, and operational best practices across engineering organizations.
  • Develop and maintain infrastructure-as-code, deployment automation, and operational tooling.
  • Support and improve CI/CD and GitOps-based deployment workflows.
  • Design observability strategies using monitoring, logging, tracing, and alerting platforms.
  • Participate in architecture reviews and provide guidance on scalability, resiliency, and operational excellence.
  • Mentor junior engineers and contribute to the technical growth of the broader engineering organization.
  • Act as a force multiplier by creating reusable solutions, self-service capabilities, and engineering standards that increase the effectiveness of multiple teams.
  • Drive adoption of AI-assisted engineering workflows and operational automation across the organization.
  • Drive engineering leverage initiatives that improve the productivity, reliability, and effectiveness of multiple engineering teams.
  • Influence the broader engineering organization through platform thinking, standardization, and operational simplification.

Requirements

  • 5+ years of experience leading reliability, platform engineering, infrastructure, or cloud operations initiatives in production environments.
  • Demonstrated experience operating and supporting large-scale production environments.
  • Demonstrated experience with Kubernetes and containerized workloads in production environments.
  • Demonstrated experience with cloud infrastructure platforms such as AWS, OCI, or GCP.
  • Demonstrated Linux systems administration and troubleshooting skills.
  • Demonstrated experience developing automation and tooling using Python, Go, Bash, or similar languages.
  • Demonstrated experience with infrastructure-as-code technologies such as Terraform.
  • Demonstrated experience designing and supporting CI/CD and GitOps workflows.
  • Demonstrated understanding of networking fundamentals including DNS, load balancing, TLS/SSL, routing, and service networking.
  • Demonstrated experience troubleshooting distributed systems and leading production incident response efforts.
  • Demonstrated track record of reducing operational complexity through automation, platform engineering, or process transformation initiatives.
  • Proven ability to influence technical decisions across teams and drive engineering improvements beyond direct ownership.
  • Ability to participate in an on-call rotation supporting production systems.

Preferred Qualifications

  • Experience with GitOps platforms such as ArgoCD.
  • Experience operating multi-region or hybrid-cloud environments.
  • Experience with observability platforms such as Prometheus, Grafana, Loki, OpenTelemetry, or similar technologies.
  • Experience designing and operating platform engineering solutions and self-service infrastructure.
  • Experience supporting high-scale SaaS environments.
  • Understanding of release strategies such as canary, blue/green, progressive delivery, and feature flag-based deployments.
  • Experience with capacity planning, performance engineering, and resilience testing.
  • Familiarity with security, compliance, and regulatory requirements in production environments.
  • Experience using AI-assisted development, automation, or operational tooling to improve engineering productivity and service reliability.
  • Experience applying AI-assisted engineering workflows to improve productivity, reliability, or operational efficiency at scale.
  • Experience designing platform engineering solutions that enable self-service and increase engineering leverage.
  • Experience mentoring engineers and leading cross-functional technical initiatives.
  • Demonstrated passion for automation, process improvement, operational excellence, and engineering scalability.
  • Strong communication, collaboration, and stakeholder management skills.

Benefits & conditions

Pulled from the full job description Paid parental leave AD&D insurance Parental leave 401(k) Health insurance Vision insurance Dental insurance, Medallia is committed to equal pay and transparency. The annual base salary range for this position is $128,500 - $190,000. Please note that the salary range information provided is a general guideline and combines all of the distinct labor markets within the US. It is uncommon for an individual to be hired at or near the top of the range for their role and compensation decisions are dependent on a variety of factors. Medallia considers factors such as (but not limited to) scope and responsibilities of the position, candidate’s work experience, candidate’s work location, education/training, key skills, internal peer equity, external market data, as well as, market and business considerations when making compensation decisions. Medallia also offers competitive health and wellness benefits, including but not limited to medical, dental, vision, 401(k), short-term and long-term disability, life and AD&D insurance, statutory leaves, paid parental leave, and paid holidays. Benefits and eligibility may vary by location and role.

About the company

Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for candidates, customers, employees, patients, and residents alike. We believe that every experience is a memory that can last a lifetime. Experiences shape the way people feel about a company. And they greatly influence how likely people are to advocate, contribute, and stay. At Medallia, we are committed to creating a world where organizations are loved by their customers and their employees. We empower exceptional people to create extraordinary experiences together. Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the infrastructure and applications that power a highly reliable global SaaS platform.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all