Manager, Site Reliability Engineering

Mastercard
O'Fallon, MO, United States
12 days ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$145,000.0 - $200,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Electronic Design Automation Monitoring of Systems Python (Programming Language) Linux System Administration Performance Tuning Reliability Engineering Prometheus Software Engineering
+6 more
Data Logging Scripting Grafana Kubernetes Terraform New Relic (SaaS)

Job description

Mastercard seeks a Manager, Site Reliability Engineering to lead a team ensuring secure, scalable, and highly available platforms. You will design and implement SRE best practices, build automation for deployment and operations, and drive observability across complex cloud-native systems. Partner with development and security teams to improve reliability, performance, and incident response. You’ll mentor engineers, champion continuous improvement, and help shape a culture of innovation, collaboration, and learning while working with cutting-edge technologies in a global environment., * Lead and mentor an SRE team supporting mission-critical platforms

  • Define and implement SRE best practices for reliability, scalability, and security
  • Design automation for deployments, configuration, and operations
  • Establish and improve monitoring, logging, and alerting for cloud-native systems
  • Drive incident management, root-cause analysis, and post-incident reviews
  • Collaborate with software engineering and security teams to improve system design
  • Optimize performance and capacity planning across services
  • Promote continuous improvement and a learning culture within the team

Requirements

  • Site Reliability Engineering (SRE)
  • Cloud platforms (AWS/Azure/GCP)
  • Kubernetes & container orchestration
  • Linux systems administration
  • CI/CD pipelines
  • Infrastructure as Code (Terraform/Cloud
  • Formation)
  • Monitoring & observability (Prometheus/Grafana/New Relic)
  • Incident management & on-call operations
  • Performance tuning & capacity planning
  • Scripting (Python/Bash)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

Videos

See all

Related articles

See all