> Markdown version of [/jobs/ext/1977244-site-reliability-engineer-iv](https://www.wearedevelopers.com/jobs/ext/1977244-site-reliability-engineer-iv). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer IV - **Company:** M&T Bank - **Location:** Buffalo, NY, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Agile Methodology, Amazon Web Services, Application Performance Management, Automation of Tests, Microsoft Azure, Bash Shell, Cloud Engineering, Continuous Integration, DevOps, Disaster Recovery, Fault Tolerance, Systems Analysis, Python (Programming Language), Windows PowerShell, Systems Development Life Cycle, Regression Testing, Reliability Engineering, Software Engineering, Data Logging, Scripting, Cloud Monitoring, Grafana, Deployment Automation, Terraform, Dynatrace - **Published:** August 7, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/site-reliability-engineer-iv-buffalo-ny-usa-58834450 ## About the Role Promote a culture of belonging and compliance with internal controls and regulatory requirements * Complete other related duties as assigned Tasks * Associate's degree with 9+ years of experience or Bachelor's degree with 7+ years of experience in systems analysis and/or application development * Expert experience in system design, reliability engineering, and production operations * Advanced proficiency in at least one programming or scripting language * Experience with observability and incident management tooling * Experience with cloud platforms (AWS or Azure) * Strong understanding of CI/CD, DevOps, and SDLC practices * Experience defining and implementing SLO/SLI frameworks * Experience in regulated environments such as financial services * Experience with Infrastructure as Code (Terraform) * Experience with monitoring/observability tools (Dynatrace, OpenTelemetry, Azure Monitor, Application Insights) * Experience with automated testing, deployment automation, and aaaaaa aaaN_ engineering practices * Knowledge of capacity planning, resiliency testing, disaster recovery, and high-availability architectures * Experience with Agile/DevOps operating models * Scripting/automation in PowerShell, Python, Bash (or similar) * Industry cloud certifications preferred Key requirements * ## Description Experteer Overview In this role you will lead platform reliability across the enterprise, acting as a SME in Site Reliability Engineering. You will drive reliability standards, observability, automation, and incident response to improve system stability and performance. You will mentor engineers and influence enterprise engineering practices while partnering with senior stakeholders. This is a scale-focused position at a financial services firm with a strong emphasis on risk management and operational excellence. Compensation / Benefits * Define and drive service reliability standards (SLOs, SLAs, SLIs, error budgets) * Design highly available, fault-tolerant architectures * Lead automation to improve reliability and operational excellence * Develop observability strategy using logging, monitoring, tracing, dashboards, and telemetry analytics * Design and maintain end-to-end monitoring solutions for application, infra, and customer experience * Analyze production telemetry to identify performance bottlenecks and risks * Lead incident management and post-incident RCA activities * Drive automation for self-healing systems and deployment/recovery workflows * Partner with development teams to build observable, scalable services across SDLC * Develop automated regression testing strategies and validate stability and performance * Create and improve IaC solutions using Terraform * Support Azure cloud environments and deployment automation * Engage in performance engineering, resilience, capacity planning, and workload optimization * Lead production readiness activities and DR/operational readiness reviews * Review architectures and roadmaps for reliability improvements * Mentor engineers on reliability, observability, cloud engineering, and automation * Prepare operational runbooks, incident playbooks, and knowledge articles * Communicate reliability metrics and remediation strategies to stakeholders * Participate in architecture reviews and leadership discussions * Promote a culture of belonging and compliance with internal controls and regulatory requirements * Complete other related duties as assigned Tasks * Associate's degree with 9+ years of experience or Bachelor's degree with 7+ years of experience in systems analysis and/or application development * Expert experience in system design, reliability engineering, and production operations * Advanced proficiency in at least one programming or scripting language * Experience with observability and incident management tooling * Experience with cloud platforms (AWS or Azure) * Strong understanding of CI/CD, DevOps, and SDLC practices * Experience defining and implementing SLO/SLI frameworks * Experience in regulated environments such as financial services * Experience with Infrastructure as Code (Terraform) * Experience with monitoring/observability tools (Dynatrace, OpenTelemetry, Azure Monitor, Application Insights) * Experience with automated testing, deployment automation, and reliability engineering practices * Knowledge of capacity planning, resiliency testing, disaster recovery, and high-availability architectures * Experience with Agile/DevOps operating models * Scripting/automation in PowerShell, Python, Bash (or similar) * Industry cloud certifications preferred Key requirements * ## Related Videos - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SMART Test Automation - the experience of legacy transformation](https://www.wearedevelopers.com/videos/100215-smart-test-automation-the-experience-of-legacy-transformation) - [The journey from developer to devops - what i've learnt along the way](https://www.wearedevelopers.com/videos/238-the-journey-from-developer-to-devops-what-i-ve-learnt-along-the-way) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)