> Markdown version of [/jobs/ext/2394356-vice-president-site-reliability-engineering](https://www.wearedevelopers.com/jobs/ext/2394356-vice-president-site-reliability-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Vice President, Site Reliability Engineering - **Company:** BNY - **Location:** London, UK - **Salary:** £97,736.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Agile Methodology, DevOps, Distributed Systems, Design of User Interfaces, Monitoring of Systems, Python (Programming Language), Machine Learning, Reliability Engineering, Ansible, Prometheus, Software Deployment, Software Engineering, Scripting, Cloud Platform System, ReactJS, Grafana, Software Troubleshooting, Backend, AngularJS, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Front End Software Development, Api Design, ArcSight Event Correlation, Splunk, Appdynamics - **Published:** August 5, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5826430221 ## About the Role + Bachelor degree in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience. + Strong full-stack development experience, with hands-on expertise in Python and Java for backend or service-layer engineering. + Strong working knowledge of front-end development using React or Angular, including building interfaces for operational or engineering use cases. + Proven experience designing and deploying end-to-end solutions, from application development through production deployment and operational support. + Experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or similar roles supporting business-critical applications. + Strong foundation in Linux/Unix systems administration, scripting, troubleshooting, and infrastructure concepts. + Hands-on experience with Ansible and Kubernetes in enterprise or production environments. + Demonstrated ability to define and operationalize SLIs, SLOs, dashboards, alerts, and health indicators. + Hands-on experience with enterprise monitoring and observability platforms including Prometheus, Grafana, AppDynamics, and Splunk. + Strong troubleshooting, analytical, and problem-solving skills in complex distributed or production environments. + Strong verbal and written communication skills, with the ability to collaborate effectively across technical and non-technical stakeholders., + Experience building centralized internal platforms or shared engineering services for operational or enterprise users. + Experience applying AIOps, machine learning, or intelligent automation within production support or reliability engineering environments. + Exposure to CI/CD pipelines, infrastructure as code, API-driven automation, and modern software delivery practices. + Experience supporting distributed systems, cloud-native platforms, or container-based architectures. + Knowledge of Agile, DevOps, and SRE operating models, including continuous improvement and blameless post-incident practices. + Ability to influence engineering standards and drive adoption of common tooling and automation patterns across teams. ## Description BNY is seeking a Vice President - Site Reliability Engineer to design, build, deploy, and scale resilient, automated, and centrally managed engineering solutions for Production Services. This role is ideal for a strong full-stack engineer who combines application development, UI engineering, backend services, infrastructure automation, and production reliability expertise. The successful candidate will build reusable platforms, internal tools, and automation capabilities that improve operational efficiency, reduce manual effort, strengthen resiliency, and enable Production Services teams to support critical business platforms more effectively. This role requires a hands-on engineer who can take solutions from concept and development through deployment, operationalization, and continuous improvement. In this role, you'll make an impact in the following ways: * Design, develop, and deploy centralized engineering solutions that improve operational efficiency, reduce toil, and enhance resiliency across Production Services. * Build full-stack applications and internal engineering tools, including backend services, APIs, automation layers, and user-facing interfaces using technologies such as Python, Java, React, or Angular. * Engineer scalable solutions that support central operational use cases such as self-service tooling, operational dashboards, alert enrichment, incident reduction, service recovery, and workflow automation. * Develop reusable frameworks and components that can be adopted broadly across Production Services teams to standardize and accelerate operational processes. * Automate infrastructure, deployment, configuration, and runtime support activities using tools such as Ansible and Kubernetes. * Define, implement, and continuously improve Service Level Indicators, Service Level Objectives, and service health measures aligned to operational and business priorities. * Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk. * Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive insights, and proactive issue prevention. * Partner with engineering, infrastructure, production support, security, and risk teams to ensure developed solutions are secure, scalable, supportable, and aligned to enterprise standards. * Identify manual, fragmented, or repetitive processes across Production Services and convert them into efficient, automated, centrally consumable solutions. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)