> Markdown version of [/jobs/ext/2805408-site-reliability-engineering-sre-lead](https://www.wearedevelopers.com/jobs/ext/2805408-site-reliability-engineering-sre-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineering (SRE ) lead - **Company:** U.S. Bank, National Association - **Location:** Atlanta, GA, United States - **Experience:** Expert - **Salary:** $111,605.0 - $131,300.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, JIRA, Microsoft Azure, Relational Databases, DevOps, Distributed Systems, Github, Python (Programming Language), Windows PowerShell, Reliability Engineering, Ansible, Prometheus, Shell Script, Software Engineering, SQL Databases, Datadog, Cloud Platform System, Cloud Monitoring, Grafana, Mttr, Infrastructure as Code (IaC), Gitlab, Kubernetes, Cloudwatch, Restful APIs, Terraform, Splunk, Dynatrace, Docker, Jenkins, Servicenow - **Published:** September 9, 2026 - **Apply:** https://usbank.wd1.myworkdayjobs.com/US_Bank_Careers/job/Atlanta-GA/Site-Reliability-Engineering--SRE---lead_2026-0025484 ## About the Role Bachelor's degree, or equivalent work experience - Six to eight years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development Preferred Skills/Experience * Strong expertise in Site Reliability Engineering (SRE), DevOps, Production Support, Platform Engineering, and Distributed Systems Operations. * Experience leading technical teams, incident response efforts, workload prioritization, and reliability improvement programs. * Advanced knowledge of Incident Management, Problem Management, Change Management, and Root Cause Analysis (RCA) methodologies. * Hands-on experience with AWS, Azure, Kubernetes, Docker, and cloud-native infrastructure platforms. * Proficiency with Python, PowerShell, Shell Scripting, and automation frameworks for operational efficiency and reliability engineering. * Experience building and supporting CI/CD pipelines using tools such as GitHub Actions, Azure DevOps, Jenkins, or GitLab. * Strong expertise in Monitoring and Observability Solutions including Datadog, Splunk, Dynatrace, Grafana, Prometheus, CloudWatch, Azure Monitor, and OpenTelemetry. * Experience with ServiceNow, Jira, Terraform, Ansible, REST APIs, SQL/Relational Databases, along with excellent stakeholder communication and leadership skills. Preferred Certifications * AWS Certified Solutions Architect, DevOps Engineer, or equivalent AWS certification * Microsoft Azure Administrator, Architect, or DevOps Engineer certification * Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD) ## Description * Lead the troubleshooting and resolution of complex production incidents, including application failures, API issues, cloud platform outages, performance degradation, and operational disruptions. * Conduct comprehensive root cause analysis (RCA), impact assessments, mitigation planning, and implementation of permanent corrective actions. * Design and enhance monitoring, observability, alerting, dashboards, health checks, and operational runbooks to improve platform reliability and availability. * Drive automation initiatives using scripting, Infrastructure as Code (IaC), CI/CD pipelines, and self-healing capabilities to reduce manual operational effort. * Partner with software engineering, infrastructure, and product teams to identify, prioritize, and remediate recurring reliability issues. * Serve as the Incident Commander during major incidents, coordinating cross-functional response teams and driving restoration activities. * Provide leadership, coaching, mentoring, and workload management for SRE, DevOps, and production support engineers. * Utilize operational metrics including MTTR, MTTD, SLA compliance, backlog health, incident volume, and problem closure rates to drive continuous improvement and operational excellence. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)