BizOps Engineer II

Mastercard
O'Fallon, MO, United States
1 day ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$125,000.0 - $165,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Computer Programming Information Engineering Monitoring of Systems Python (Programming Language) Linux System Administration Performance Tuning Reliability Engineering Prometheus
+10 more
SQL Databases Scripting System Availability Grafana Reliability of Systems Containerization Kubernetes Terraform Splunk Docker

Job description

Mastercard is seeking a BizOps Engineer II to join our Information Technology & Data Management team in Financial Services. In this role, you will optimize and support mission-critical platforms that power secure, high-volume transactions worldwide. You’ll design and implement robust monitoring, automation, and incident management solutions to ensure high availability, performance, and scalability. Collaborating closely with software engineers, data engineers, and product teams, you will troubleshoot complex issues, analyze system metrics, and drive continuous improvement across infrastructure and applications. This position offers the opportunity to work with cutting-edge cloud and data technologies while influencing operational best practices and reliability standards. You’ll contribute to incident response, root-cause analysis, and post-incident reviews, helping to build more resilient systems. Mastercard’s culture emphasizes innovation, collaboration, and continuous learning, giving you room to experiment with new tools and approaches. If you are passionate about system reliability, automation, and bridging the gap between development and operations in a dynamic, global environment, this role provides a chance to make a tangible impact on secure digital payments worldwide.

Responsibilities

  • Design, implement, and maintain monitoring, alerting, and observability for mission-critical applications and infrastructure.
  • Automate operational tasks, deployments, and remediation workflows to improve reliability and reduce manual intervention.
  • Collaborate with software and data engineering teams to optimize system performance, scalability, and resilience.
  • Lead and participate in incident response, troubleshooting complex production issues, and driving timely resolution.
  • Conduct root-cause analysis and implement long-term fixes to prevent recurrence of incidents.
  • Support CI/CD pipelines and release processes to enable safe, rapid, and reliable deployments.
  • Analyze system metrics and logs to identify bottlenecks, trends, and optimization opportunities.
  • Contribute to reliability best practices, runbooks, and operational documentation.
  • Partner with security and compliance teams to ensure systems meet regulatory and security requirements.
  • Mentor junior team members and help foster a culture of continuous improvement and learning.

Requirements

  • Site Reliability Engineering (SRE) practices
  • Cloud platforms (AWS, GCP, or Azure; AWS preferred)
  • Infrastructure as Code (Terraform, Cloud
  • Formation, or similar)
  • Linux systems administration
  • Containerization and orchestration (Docker, Kubernetes)
  • CI/CD pipelines (Jenkins, Git
  • Lab CI, or similar)
  • Monitoring and observability (Splunk, Prometheus, Grafana, Cloud
  • Watch)
  • Scripting/programming (Python, Bash, or similar)
  • SQL and basic data querying/analysis
  • Incident management and root-cause analysis

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all