Site Reliability Engineering (SRE) Manager

M&T Bank
Buffalo, NY, United States
12 days ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$139,700.0 - $232,900.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Applications Architecture Microsoft Azure Bash Shell Cloud Computing Cloud Engineering Continuous Integration DevOps Disaster Recovery Distributed Systems Python (Programming Language)
+18 more
Knowledge Management Windows PowerShell Reliability Engineering Cloud Services Software Engineering Datadog Data Logging Cloud Monitoring System Availability Grafana Reliability of Systems Infrastructure as Code (IaC) Infrastructure Automation Frameworks Information Technology Performance Monitor Splunk Dynatrace Microservices

Job description

Reliability & Operational Excellence

  • Define and execute SRE strategies that improve system reliability, availability, scalability, and performance.
  • Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational health metrics.
  • Lead production readiness reviews, disaster recovery testing, resilience assessments, and operational risk mitigation activities.
  • Drive continuous improvement of application stability, service availability, and customer experience.

Incident & Problem Management

  • Lead major incident response and escalation management for critical production issues.
  • Oversee root cause analysis (RCA) processes and ensure corrective actions are implemented and tracked to completion.
  • Drive reduction of recurring incidents through engineering improvements, automation, and proactive monitoring.
  • Provide executive-level communication during significant incidents and service disruptions.

Observability & Automation

  • Establish monitoring, alerting, logging, tracing, and observability standards across supported platforms.
  • Lead implementation of dashboards and operational metrics that provide visibility into service health and customer impact.
  • Drive automation initiatives that reduce manual operational effort, improve recovery times, and increase engineering efficiency.
  • Promote Infrastructure as Code (IaC), CI/CD integration, automated remediation, and self-service operational capabilities.

Cloud & Platform Reliability

  • Partner with Engineering and Infrastructure teams to support cloud-native and hybrid application environments.
  • Ensure applications are designed and operated using resilient, scalable, and supportable architectures.
  • Support modernization initiatives involving Azure cloud services, containers, APIs, microservices, and platform engineering practices.
  • Evaluate vendor platforms and third-party services to ensure reliability and operational readiness.

AI & Modern Operations

  • Drive adoption of AI and Generative AI capabilities to improve incident response, troubleshooting, observability, and operational efficiency.
  • Identify opportunities for intelligent automation, anomaly detection, automated diagnostics, and AI-assisted knowledge management.
  • Promote responsible AI adoption aligned with enterprise security, governance, and risk standards.

People Leadership

  • Recruit, develop, coach, and retain high-performing Site Reliability Engineers, Production Engineers, Automation Engineers, and Observability Engineers.
  • Establish career paths, skill development plans, and succession strategies.
  • Foster a culture of ownership, accountability, innovation, collaboration, and continuous learning.
  • Manage staffing, performance management, compensation recommendations, and organizational development activities.

Risk & Governance

  • Ensure adherence to enterprise risk, cybersecurity, regulatory, and operational control standards.
  • Identify and escalate operational risks impacting critical services or customer experiences.
  • Support audits, regulatory reviews, disaster recovery exercises, and operational governance programs.

Scope of Responsibilities

Leads teams responsible for:

  • Site Reliability Engineering (SRE)
  • Production Support
  • Observability Engineering
  • Incident Management
  • Operational Automation
  • Cloud Reliability
  • Platform Operations

Responsible for reliability and operational health across multiple applications, platforms, cloud services, and vendor-supported solutions.

Supervisory Responsibilities

Typically manages 10-20 direct and indirect reports including SRE Engineers, Production Engineers, Technical Leads, and Engineering Managers.

Requirements

  • 10+ years of technology experience with application support, infrastructure, cloud, software engineering, or reliability engineering responsibilities.
  • 5+ years of leadership experience managing engineering, operations, or SRE teams.
  • Experience managing production systems supporting critical business functions.
  • Strong knowledge of Site Reliability Engineering principles, including SLOs, observability, automation, incident management, and operational excellence.
  • Experience leading major incident response, root cause analysis, and service restoration efforts.
  • Experience with cloud platforms, distributed systems, APIs, and modern application architectures.
  • Strong communication, analytical, decision-making, and stakeholder management skills., * Bachelor’s degree in Computer Science, Engineering, Information Technology, or related field.
  • Experience leading SRE or Production Engineering organizations.
  • Experience with Azure cloud technologies and cloud-native architectures.
  • Experience with observability platforms such as Dynatrace, Splunk, Datadog, Grafana, Azure Monitor, or OpenTelemetry.
  • Experience with scripting and automation technologies including PowerShell, Python, Bash, and APIs.
  • Experience with CI/CD, Infrastructure as Code, DevOps, and Platform Engineering practices.
  • Experience implementing operational AI use cases including incident analysis, observability analytics, and automated diagnostics.
  • Financial services or other highly regulated industry experience preferred.

Benefits & conditions

M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

About the company

M&T Bank Corporation is an Equal Opportunity/Affirmative Action Employer, including disabilities and veterans.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:27 min

Defining DevOps through its historical origins and foundational texts

Sonal Patil · LIVE

Videos

See all

Related articles

See all