Site Reliability Engineer

Mastercard
O'Fallon, MO, United States
3 days ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$135,000.0 - $180,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Configuration Management Computer Programming Customer Data Management Distributed Systems Monitoring of Systems Python (Programming Language) Linux System Administration Performance Tuning
+16 more
Reliability Engineering Prometheus Datadog Data Logging Scripting Grafana Git Containerization Kubernetes Infrastructure Automation Frameworks Terraform New Relic (SaaS) Docker Jenkins Golang Microservices

Job description

Mastercard is seeking a Senior Site Reliability Engineer to enhance reliability, scalability, and performance of our critical IT & Data Management platforms. You will design resilient architectures, automate deployments, and champion observability to ensure always-on services. Collaborating with cross-functional teams, you’ll identify and resolve production issues, implement robust incident management, and drive SRE best practices. This role offers the opportunity to work with cutting-edge cloud and container technologies in a culture that values innovation, collaboration, and continuous growth., * Design and maintain highly available, scalable, and secure cloud infrastructure for Mastercard’s IT & Data Management platforms.

  • Build and improve automation for deployments, configuration management, and infrastructure provisioning.
  • Implement and refine monitoring, logging, and alerting to ensure service reliability and rapid incident detection.
  • Lead and participate in incident response, root cause analysis, and post-incident reviews to drive continuous improvement.
  • Partner with development and data teams to embed SRE best practices, including SLIs/SLOs and capacity planning.
  • Optimize system performance and cost efficiency across distributed, cloud-native environments.
  • Develop tools and scripts to reduce toil and improve operational excellence.
  • Contribute to security, compliance, and governance standards within production environments.

Requirements

  • Site Reliability Engineering (SRE)
  • Cloud platforms (AWS, Azure, or GCP)
  • Kubernetes and containerization (Docker)
  • Infrastructure as Code (Terraform/Cloud
  • Formation)
  • CI/CD pipelines (Jenkins, Git
  • Hub Actions, Git
  • Lab CI)
  • Linux systems administration
  • Monitoring and observability (Prometheus, Grafana, Datadog, New Relic)
  • Scripting/programming (Python, Go, Bash)
  • Distributed systems and microservices
  • Incident management and on-call operations

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all