Site Reliability Engineer IV

M&T Bank
Buffalo, NY, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours

Tech stack

Agile Methodology Amazon Web Services Application Performance Management Automation of Tests Microsoft Azure Bash Shell Cloud Engineering Continuous Integration DevOps Disaster Recovery Fault Tolerance Systems Analysis
+13 more
Python (Programming Language) Windows PowerShell Systems Development Life Cycle Regression Testing Reliability Engineering Software Engineering Data Logging Scripting Cloud Monitoring Grafana Deployment Automation Terraform Dynatrace

Job description

Experteer Overview In this role you will lead platform reliability across the enterprise, acting as a SME in Site Reliability Engineering. You will drive reliability standards, observability, automation, and incident response to improve system stability and performance. You will mentor engineers and influence enterprise engineering practices while partnering with senior stakeholders. This is a scale-focused position at a financial services firm with a strong emphasis on risk management and operational excellence. Compensation / Benefits * Define and drive service reliability standards (SLOs, SLAs, SLIs, error budgets) * Design highly available, fault-tolerant architectures * Lead automation to improve reliability and operational excellence * Develop observability strategy using logging, monitoring, tracing, dashboards, and telemetry analytics * Design and maintain end-to-end monitoring solutions for application, infra, and customer experience * Analyze production telemetry to identify performance bottlenecks and risks * Lead incident management and post-incident RCA activities * Drive automation for self-healing systems and deployment/recovery workflows * Partner with development teams to build observable, scalable services across SDLC * Develop automated regression testing strategies and validate stability and performance * Create and improve IaC solutions using Terraform * Support Azure cloud environments and deployment automation * Engage in performance engineering, resilience, capacity planning, and workload optimization * Lead production readiness activities and DR/operational readiness reviews * Review architectures and roadmaps for reliability improvements * Mentor engineers on reliability, observability, cloud engineering, and automation * Prepare operational runbooks, incident playbooks, and knowledge articles * Communicate reliability metrics and remediation strategies to stakeholders * Participate in architecture reviews and leadership discussions * Promote a culture of belonging and compliance with internal controls and regulatory requirements * Complete other related duties as assigned Tasks * Associate’s degree with 9+ years of experience or Bachelor’s degree with 7+ years of experience in systems analysis and/or application development * Expert experience in system design, reliability engineering, and production operations * Advanced proficiency in at least one programming or scripting language * Experience with observability and incident management tooling * Experience with cloud platforms (AWS or Azure) * Strong understanding of CI/CD, DevOps, and SDLC practices * Experience defining and implementing SLO/SLI frameworks * Experience in regulated environments such as financial services * Experience with Infrastructure as Code (Terraform) * Experience with monitoring/observability tools (Dynatrace, OpenTelemetry, Azure Monitor, Application Insights) * Experience with automated testing, deployment automation, and reliability engineering practices * Knowledge of capacity planning, resiliency testing, disaster recovery, and high-availability architectures * Experience with Agile/DevOps operating models * Scripting/automation in PowerShell, Python, Bash (or similar) * Industry cloud certifications preferred Key requirements *

Requirements

Promote a culture of belonging and compliance with internal controls and regulatory requirements * Complete other related duties as assigned Tasks * Associate’s degree with 9+ years of experience or Bachelor’s degree with 7+ years of experience in systems analysis and/or application development * Expert experience in system design, reliability engineering, and production operations * Advanced proficiency in at least one programming or scripting language * Experience with observability and incident management tooling * Experience with cloud platforms (AWS or Azure) * Strong understanding of CI/CD, DevOps, and SDLC practices * Experience defining and implementing SLO/SLI frameworks * Experience in regulated environments such as financial services * Experience with Infrastructure as Code (Terraform) * Experience with monitoring/observability tools (Dynatrace, OpenTelemetry, Azure Monitor, Application Insights) * Experience with automated testing, deployment automation, and aaaaaa aaaN_ engineering practices * Knowledge of capacity planning, resiliency testing, disaster recovery, and high-availability architectures * Experience with Agile/DevOps operating models * Scripting/automation in PowerShell, Python, Bash (or similar) * Industry cloud certifications preferred Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Scaling IT operations for a major finance cloud

Linda Linda +1 · WWC 2024

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · WWC 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:06 min

Developer experience and project variety at scale

Alexandra Petri · WWC 2023

12:08 min

Comparing Keptn orchestration capabilities against alternative software operators

Thomas Schütz · LIVE

Videos

See all

Related articles

See all