> Markdown version of [/jobs/ext/2293836-workforce-compute-ops-sre-lead-infrastructure](https://www.wearedevelopers.com/jobs/ext/2293836-workforce-compute-ops-sre-lead-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Workforce Compute Ops SRE Lead Infrastructure... - **Company:** Wells Fargo - **Location:** Woodbridge Township, NJ, United States - **Experience:** Expert - **Salary:** $119,000.0 - $224,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Data Analysis, Business Software, Software Debugging, Monitoring of Systems, Python (Programming Language), Reliability Engineering, Prometheus, Software Engineering, Software Systems, Data Processing, Large Language Models, Grafana, Prompt Engineering, Generative AI, Kubernetes, Infrastructure Automation Frameworks, Data Analytics, Performance Monitor, Low-code, Virtual Agents, Api Design, Splunk - **Published:** August 29, 2026 - **Apply:** https://www.juju.com/job/00000000gpg75o ## About the Role + 5+ years of Technology Infrastructure Engineering and Solutions experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education + 5+ years of experience developing software, automation, or data processing solutions using Python + 5+ years of experience with observability and monitoring technologies such as Splunk, Grafana, Prometheus, Elastic, or similar tools Desired Qualifications: + Experience providing technical leadership and mentoring engineers + Strong knowledge of Site Reliability Engineering principles and practices, including observability, service level indicators (SLIs), service level objectives (SLOs), error budgets, incident management, problem management, and operational resilience + Experience developing Generative AI or Agentic AI solutions + Knowledge of large language models (LLMs), prompt engineering, retrieval-augmented generation (RAG), AI-assisted workflows, and API-based AI integrations + Experience with workflow automation and low-code/no-code platforms + Experience using operational telemetry and data analytics to identify trends, anomalies, and opportunities for proactive remediation + Experience with observability and monitoring platforms such as Splunk, Grafana, Prometheus, Elastic, or similar technologies + Experience with containers, Kubernetes, and Infrastructure as Code practices + Strong problem-solving, communication, and collaboration skills + Ability to provide technical direction and influence engineering practices across teams ## Description Wells Fargo is seeking a Lead Site Reliability Engineer to provide technical leadership, mentorship, and guidance to a team responsible for the stability, reliability, availability, performance, and continuous improvement of enterprise workplace technology platforms. In this role, you will establish and promote Site Reliability Engineering (SRE) standards and best practices across observability, service level indicators (SLIs), service level objectives (SLOs), error budgets, incident and problem management, automation, and operational resilience. You will leverage telemetry and data-driven insights to identify risks, reduce incident frequency and recurrence, and drive continuous reliability improvements. You will also help improve the stability of workplace technology platforms through intelligent automation, Agentic AI capabilities, and low-code/no-code solutions. This includes providing technical leadership for AI-assisted incident triage, root cause analysis, workflow automation, proactive remediation, and self-healing capabilities to improve operational efficiency and the end-user experience. In this role, you will: + Lead complex initiatives to develop infrastructure solutions that support business applications. + Participate in projects intended to improve, modernize, and enhance technology infrastructure. + Evaluate internal and external software solutions to support target-state architecture objectives. + Review and analyze high-impact outages and implement processes to reduce future operational risk. + Design, build, deploy, and maintain infrastructure solutions in collaboration with technology teams and third-party vendors. + Design, code, test, debug, and document solutions using Agile development practices. + Influence technical designs and implementation plans while identifying project risks and resource requirements. + Provide technical leadership and guidance to engineers and partners across the organization. + Direct risk and control activities by ensuring adherence to policies, procedures, and operational standards. + Recommend solutions that improve efficiency, manage costs, and achieve business objectives. + Collaborate with peers, leaders, customers, and vendors to resolve issues and deliver technology solutions., Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit's risk appetite and all risk and compliance program requirements. ## Related Videos - [Reimagining app development with Low-code and AI](https://www.wearedevelopers.com/videos/1651-reimagining-app-development-with-low-code-and-ai) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [What If Apps Built Themselves? AI-Powered Low-Code for the Industrial Enterprise](https://www.wearedevelopers.com/videos/2079-what-if-apps-built-themselves-ai-powered-low-code-for-the-industrial-enterprise) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How Much Does a Software Engineer Make? Realistic Software Engineering Salaries](https://www.wearedevelopers.com/magazine/425-how-much-does-a-software-engineer-make-realistic-software-engineering-salaries) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)