Lead Site Reliability Engineer

Mastercard
O'Fallon, MO, United States
1 day ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$155,000.0 - $205,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Cloud Computing Cyber Security Computer Programming Continuous Integration Linux Distributed Systems Monitoring of Systems Python (Programming Language) Reliability Engineering Prometheus
+10 more
Datadog Data Logging Scripting Cloud Platform System Grafana Git Kubernetes Terraform Splunk Jenkins

Job description

Mastercard is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and security for mission-critical financial services platforms. You will design and optimize cloud-native, highly available systems, implement SRE best practices, and lead incident response and postmortems. Partnering with IT and Cybersecurity teams, you’ll automate deployments, observability, and resilience testing while mentoring engineers. Ideal candidates bring deep experience with cloud, CI/CD, infrastructure-as-code, and securing large-scale, distributed systems in a regulated environment., * Lead design and operation of highly available, secure, and scalable financial services platforms.

  • Define and implement SRE best practices, including SLOs, SLIs, and error budgets.
  • Architect and maintain cloud-native infrastructure using infrastructure-as-code and automation.
  • Own incident response, root cause analysis, and postmortems for critical production issues.
  • Drive observability across systems with robust monitoring, logging, and alerting solutions.
  • Collaborate closely with IT and Cybersecurity teams to embed security and compliance into the stack.
  • Optimize performance, capacity planning, and cost management for large-scale distributed systems.
  • Mentor and guide engineers on SRE principles, tooling, and operational excellence.
  • Continuously improve CI/CD pipelines and deployment strategies for safer, faster releases.
  • Champion a culture of reliability, innovation, and continuous improvement within the team.

Requirements

  • Site Reliability Engineering (SRE)
  • Public cloud platforms (AWS, GCP, or Azure)
  • Kubernetes and container orchestration
  • Linux systems engineering and administration
  • Infrastructure as Code (Terraform, Cloud
  • Formation, or similar)
  • CI/CD pipelines (Jenkins, Git
  • Lab CI, Git
  • Hub Actions, or similar)
  • Monitoring and observability (Prometheus, Grafana, Datadog, Splunk, etc.)
  • Scripting/programming (Python, Go, or similar)
  • Security and compliance for financial/regulated environments
  • Incident management and on-call operations

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all