> Markdown version of [/jobs/ext/2922158-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2922158-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** Mastercard - **Location:** O'Fallon, MO, United States - **Experience:** Expert - **Salary:** $155,000.0 - $205,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Cloud Computing, Cyber Security, Computer Programming, Continuous Integration, Linux, Distributed Systems, Monitoring of Systems, Python (Programming Language), Reliability Engineering, Prometheus, Datadog, Data Logging, Scripting, Cloud Platform System, Grafana, Git, Kubernetes, Terraform, Splunk, Jenkins - **Published:** September 15, 2026 - **Apply:** https://find.jobs/jobs-near-me/apply/ats-redirect/?id=2968576607-2 ## About the Role * Site Reliability Engineering (SRE) * Public cloud platforms (AWS, GCP, or Azure) * Kubernetes and container orchestration * Linux systems engineering and administration * Infrastructure as Code (Terraform, Cloud * Formation, or similar) * CI/CD pipelines (Jenkins, Git * Lab CI, Git * Hub Actions, or similar) * Monitoring and observability (Prometheus, Grafana, Datadog, Splunk, etc.) * Scripting/programming (Python, Go, or similar) * Security and compliance for financial/regulated environments * Incident management and on-call operations ## Description Mastercard is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and security for mission-critical financial services platforms. You will design and optimize cloud-native, highly available systems, implement SRE best practices, and lead incident response and postmortems. Partnering with IT and Cybersecurity teams, you'll automate deployments, observability, and resilience testing while mentoring engineers. Ideal candidates bring deep experience with cloud, CI/CD, infrastructure-as-code, and securing large-scale, distributed systems in a regulated environment., * Lead design and operation of highly available, secure, and scalable financial services platforms. * Define and implement SRE best practices, including SLOs, SLIs, and error budgets. * Architect and maintain cloud-native infrastructure using infrastructure-as-code and automation. * Own incident response, root cause analysis, and postmortems for critical production issues. * Drive observability across systems with robust monitoring, logging, and alerting solutions. * Collaborate closely with IT and Cybersecurity teams to embed security and compliance into the stack. * Optimize performance, capacity planning, and cost management for large-scale distributed systems. * Mentor and guide engineers on SRE principles, tooling, and operational excellence. * Continuously improve CI/CD pipelines and deployment strategies for safer, faster releases. * Champion a culture of reliability, innovation, and continuous improvement within the team. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)