> Markdown version of [/jobs/ext/291740-lead-infrastructure-engineer-sre](https://www.wearedevelopers.com/jobs/ext/291740-lead-infrastructure-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Infrastructure Engineer (SRE) - **Company:** Wells Fargo - **Location:** Charlotte, NC, United States - **Experience:** Expert - **Salary:** $119,000.0 - $224,000.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Vbscript, Agile Methodology, Data Analysis, Application Performance Management, Systems Engineering, Confluence, JIRA, Bash Shell, Databases, Continuous Integration, Data Integration, Noise Reduction, Distributed Systems, Apache JMeter, Python (Programming Language), Windows PowerShell, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Prometheus, Systems Integration, Data Logging, Scripting, Grafana, Reliability of Systems, Containerization, Blazemeter, Kubernetes, Infrastructure Automation Frameworks, Performance Monitor, ArcSight Event Correlation, Splunk, Appdynamics, Dynatrace, Docker, Programming Languages - **Published:** May 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a2f5490c8d020bf8 ## About the Role Do you have experience in System performance monitoring?, * 5+ years of Technology Infrastructure Engineering and Solutions experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education * 5+ years of experience using Observability Tools with hands-on implementation of monitoring, logging, or tracing solutions utilizing Grafana, ThousandEyes, Prometheus, AppDynamics, or Splunk * 3+ years of application production support experience in complex, high-availability environments * 2+ years of experience with Confluence or Jira, * Experienced with Site Reliability Engineering (SRE) including SLO/SLI frameworks, error budgets, toil reduction, and production reliability engineering practices * Experience with database logging and monitoring concepts experience * Experience with Application performance monitoring and optimization using BlazeMeter, JMeter, Splunk, AppDynamics, or similar observability platforms * Experience with scripting or programming languages such as Bash, PowerShell, Python, Shell, VBScript, or JavaScript for automation and reliability engineering use cases * Experience and understanding of AIOps and related tools such as MoogSoft or Big Panda, including event correlation and noise reduction * Experience with one or more automation tools such as Ansible or similar infrastructure-as-code/configuration management tools * Experience with Container technologies: Kubernetes, Docker, PKS, with focus on observability and reliability patterns in distributed systems ## Description As a Lead SRE, you will be part of a high-impact team responsible for advancing and embedding SRE practices across multiple applications and critical customer journeys within the Banking Operations platform. You will play a central role in transforming how reliability, scalability, and observability are engineered and sustained-helping to shape a modern, resilient, and data-driven technology ecosystem. This team is at the forefront of driving technology transformation across the enterprise by adopting SRE-aligned capabilities, launching new tooling, automating complex operational challenges, and integrating with modern platforms and pipelines. Leveraging your background in software and systems engineering, you will ensure that onboarded applications are highly available, resilient, and fully instrumented with end-to-end observability. In this role, you will lead the adoption and evolution of observability practices-including metrics, logging, tracing, and telemetry-while promoting operational excellence through code, automation, and continuous improvement. You will introduce and scale data-driven insights, enabling smarter decision-making and proactive issue resolution across the ecosystem. You will also partner closely with application and platform engineering teams to ensure services are reliable, measurable, and continuously improving. Your work will include building and enhancing CI/CD integration, validating system reliability through rigorous testing, and driving the modernization of operational practices across the organization. In this role, you will: * Drive and lead Site Reliability Engineering capabilities at Wells Fargo Banking Operations igniting the practice, principles, and culture, leading by example. Mentor and coach engineers while scaling the SRE practice within Banking Operations and partnering with peer platform embedded SRE teams * Leverage enterprise capabilities, tools, and innovation to improve availability in a complex ecosystem by maturing observability practices including monitoring, logging, distributed tracing, synthetic monitoring, and chaos engineering with a focus on actionable insights and proactive detection * Lead the evolution of our environment introducing self-healing and autonomic capabilities, solving complex operational and systemic issues with precision including building and training models, automating cognitive processes, and leveraging telemetry to improve availability and reliability of products we provide to customers * Own and automate key SRE metrics and IT Service Operations processes including customer impact, golden signals and critical user journeys, % availability of critical business flows, SLO/SLI definition and adherence, error budget management, and real-time observability dashboards; automate incident response processes through data integration with unified communications and alerting/notification systems * Provide leadership in support responsibilities for critical applications and customer journeys onboarded to SRE including rapid remediation of issues through Agile practices, conducting blameless post mortems, driving root cause analysis, and implementing durable solutions through continuous improvement with the goal of eliminating repeat incidents ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026)