Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+16 more
Job description
Mastercard is seeking a Senior Site Reliability Engineer to enhance reliability, scalability, and performance of our critical IT & Data Management platforms. You will design resilient architectures, automate deployments, and champion observability to ensure always-on services. Collaborating with cross-functional teams, you’ll identify and resolve production issues, implement robust incident management, and drive SRE best practices. This role offers the opportunity to work with cutting-edge cloud and container technologies in a culture that values innovation, collaboration, and continuous growth., * Design and maintain highly available, scalable, and secure cloud infrastructure for Mastercard’s IT & Data Management platforms.
- Build and improve automation for deployments, configuration management, and infrastructure provisioning.
- Implement and refine monitoring, logging, and alerting to ensure service reliability and rapid incident detection.
- Lead and participate in incident response, root cause analysis, and post-incident reviews to drive continuous improvement.
- Partner with development and data teams to embed SRE best practices, including SLIs/SLOs and capacity planning.
- Optimize system performance and cost efficiency across distributed, cloud-native environments.
- Develop tools and scripts to reduce toil and improve operational excellence.
- Contribute to security, compliance, and governance standards within production environments.
Requirements
- Site Reliability Engineering (SRE)
- Cloud platforms (AWS, Azure, or GCP)
- Kubernetes and containerization (Docker)
- Infrastructure as Code (Terraform/Cloud
- Formation)
- CI/CD pipelines (Jenkins, Git
- Hub Actions, Git
- Lab CI)
- Linux systems administration
- Monitoring and observability (Prometheus, Grafana, Datadog, New Relic)
- Scripting/programming (Python, Go, Bash)
- Distributed systems and microservices
- Incident management and on-call operations
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Is Software Engineering Over-Saturated?
Why Upskilling And Reskilling is Important For Developers
Highest Paying Tech Companies for Developers