Site Reliability Engineering (SRE) Manager
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+18 more
Job description
Reliability & Operational Excellence
- Define and execute SRE strategies that improve system reliability, availability, scalability, and performance.
- Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational health metrics.
- Lead production readiness reviews, disaster recovery testing, resilience assessments, and operational risk mitigation activities.
- Drive continuous improvement of application stability, service availability, and customer experience.
Incident & Problem Management
- Lead major incident response and escalation management for critical production issues.
- Oversee root cause analysis (RCA) processes and ensure corrective actions are implemented and tracked to completion.
- Drive reduction of recurring incidents through engineering improvements, automation, and proactive monitoring.
- Provide executive-level communication during significant incidents and service disruptions.
Observability & Automation
- Establish monitoring, alerting, logging, tracing, and observability standards across supported platforms.
- Lead implementation of dashboards and operational metrics that provide visibility into service health and customer impact.
- Drive automation initiatives that reduce manual operational effort, improve recovery times, and increase engineering efficiency.
- Promote Infrastructure as Code (IaC), CI/CD integration, automated remediation, and self-service operational capabilities.
Cloud & Platform Reliability
- Partner with Engineering and Infrastructure teams to support cloud-native and hybrid application environments.
- Ensure applications are designed and operated using resilient, scalable, and supportable architectures.
- Support modernization initiatives involving Azure cloud services, containers, APIs, microservices, and platform engineering practices.
- Evaluate vendor platforms and third-party services to ensure reliability and operational readiness.
AI & Modern Operations
- Drive adoption of AI and Generative AI capabilities to improve incident response, troubleshooting, observability, and operational efficiency.
- Identify opportunities for intelligent automation, anomaly detection, automated diagnostics, and AI-assisted knowledge management.
- Promote responsible AI adoption aligned with enterprise security, governance, and risk standards.
People Leadership
- Recruit, develop, coach, and retain high-performing Site Reliability Engineers, Production Engineers, Automation Engineers, and Observability Engineers.
- Establish career paths, skill development plans, and succession strategies.
- Foster a culture of ownership, accountability, innovation, collaboration, and continuous learning.
- Manage staffing, performance management, compensation recommendations, and organizational development activities.
Risk & Governance
- Ensure adherence to enterprise risk, cybersecurity, regulatory, and operational control standards.
- Identify and escalate operational risks impacting critical services or customer experiences.
- Support audits, regulatory reviews, disaster recovery exercises, and operational governance programs.
Scope of Responsibilities
Leads teams responsible for:
- Site Reliability Engineering (SRE)
- Production Support
- Observability Engineering
- Incident Management
- Operational Automation
- Cloud Reliability
- Platform Operations
Responsible for reliability and operational health across multiple applications, platforms, cloud services, and vendor-supported solutions.
Supervisory Responsibilities
Typically manages 10-20 direct and indirect reports including SRE Engineers, Production Engineers, Technical Leads, and Engineering Managers.
Requirements
- 10+ years of technology experience with application support, infrastructure, cloud, software engineering, or reliability engineering responsibilities.
- 5+ years of leadership experience managing engineering, operations, or SRE teams.
- Experience managing production systems supporting critical business functions.
- Strong knowledge of Site Reliability Engineering principles, including SLOs, observability, automation, incident management, and operational excellence.
- Experience leading major incident response, root cause analysis, and service restoration efforts.
- Experience with cloud platforms, distributed systems, APIs, and modern application architectures.
- Strong communication, analytical, decision-making, and stakeholder management skills., * Bachelor’s degree in Computer Science, Engineering, Information Technology, or related field.
- Experience leading SRE or Production Engineering organizations.
- Experience with Azure cloud technologies and cloud-native architectures.
- Experience with observability platforms such as Dynatrace, Splunk, Datadog, Grafana, Azure Monitor, or OpenTelemetry.
- Experience with scripting and automation technologies including PowerShell, Python, Bash, and APIs.
- Experience with CI/CD, Infrastructure as Code, DevOps, and Platform Engineering practices.
- Experience implementing operational AI use cases including incident analysis, observability analytics, and automated diagnostics.
- Financial services or other highly regulated industry experience preferred.
Benefits & conditions
M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.
About the company
M&T Bank Corporation is an Equal Opportunity/Affirmative Action Employer, including disabilities and veterans.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Fully Remote Software Engineer Jobs
What is Software Engineering?
From developer to manager – what does it take to become an engineering manager?