> Markdown version of [/jobs/ext/2063103-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2063103-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Site Reliability Engineer - **Company:** NICE - **Location:** Southampton, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Microsoft Azure, C Sharp (Programming Language), Cloud Computing, Cloud Engineering, Cyber Security, Computer Programming, Databases, Continuous Integration, DevOps, Elasticsearch, JSON, Python (Programming Language), Microsoft SQL Server, Performance Tuning, Windows PowerShell, Reliability Engineering, Cloud Services, Prometheus, Azure DevOps Pipelines, Extensible Markup Language (XML), Data Processing, Scripting, Cloud Platform System, Cloud Monitoring, Grafana, Git, Containerization, Kubernetes, Information Technology, Bicep, Terraform, Software Version Control, Microservices - **Published:** August 15, 2026 - **Apply:** https://www.totaljobs.com/job/lead-site-reliability-engineer/nice-job107844074 ## About the Role * Must have 6+ years of experience in Site Reliability Engineering * Excellent technical, analytical and troubleshooting skills * Experience and in-depth knowledge of databases and data handling (MS-SQL, Elasticsearch, YML, JSON, XML) * Experience with Azure cloud * Significant experience in programming or advanced scripting (Python, PowerShell, C# etc.) * Experience with infrastructure/configuration as code and version control (ARM, BICEP, Git) * Strong Experience managing monitoring, alerting and dashboarding platforms (Azure Monitor, Prometheus, Grafana, Elasticsearch) * Demonstrable experience of supporting live cloud services and platforms * Expert in developing queries for dashboards and alerting for microservices. * Expertise in developing custom metrics for microservices * Collaborate with DevOps and engineering teams to establish and enforce SLOs, SLAs, and error budgets. * Production experience with Kubernetes and containerization (AKS) * Exposure to Azure DevOps pipelines is desirable (CI/CD) * Strong experience in infrastructure as a code, design and implementation strategies. * Experience with AI (tools) to automate and accelerate is a plus. * Efficient, effective, and respectful communication skills both with customers and within internal departments. Including, + Good listener, able to identify and validate assumptions. + Able to use effective questioning to confirm understanding of a customer problem and then provide help to solve it. + Methodical troubleshooting, technical skill and attention to detail used in diagnosing problems and reproducing issues in a local environment. + Multi-tasking and time-management to prioritise and switch between varied tasks. * Significant experience in platform engineering, observability, and provisioning. * Proven ability to develop and implement a strategic vision for platform services, observability, and provisioning. * Strong understanding of cyber security principles, governance, and compliance frameworks. * Strong understanding and experience of cloud platforms, containerisation, and microservices architecture. * Broad background across information technology with the ability to communicate clearly with non-security technical SMEs at a comfortable level. * Strong proficiency in technical scoping, architecture design, and integration of security tools and processes. * Ability to translate business needs into scalable, user-centric cloud solutions. * Excellent communication and collaboration skills, with a focus thought leadership and solution development. * Experience in both operational and transformation roles or a clear working understanding of both perspectives. * Knowledge of compliance with relevant frameworks, including ISO 27001, Cyber Essentials + or FEDRAMP Tooling * Kubernetes (Ideally AKS) * Azure Devops Pipelines * Elasticsearch and Cloud Observability Stacks * IAC (Bicep/Terraform) * Powershell / C# ## Description We are currently expanding our Cloud Platform Engineering team to ensure we continue to offer exemplary service to our customers. This is a very hands-on role. You will be involved in ensuring our cloud platforms are observable, measurable, reliable, scalable, and maintainable. It's likely that the successful candidate will have significant experience in a DevOps, SRE, Cloud Engineer, or Cloud Development role. NOTE - The successful candidate must have lived in the UK for 5 years and be eligible to obtain NPPV3 + Security Clearance. How will you make an impact? * Act as part of a team of SRE's that act as the 'gatekeepers' of production and actively manage the work backlog and develop reliability improvements. * Lead investigations into root cause outages, performance, and cost issues. * Lead initiatives to develop the automation of low-value tasks balanced against project delivery demands. * You will provide technical leadership and to wider Cloud Operations and Support teams along with providing oversight to the products and services they support. * Collaborate with DevOps and engineering teams to establish and enforce SLOs, SLAs, and error budgets * Develop and configure monitoring dashboards and alerts in tools like Grafana and Azure Monitor. * Installation and configuration of Observability Platform including tools like Grafana, Prometheus, Azure Monitor, Open telemetry etc. * Developing bicep modules for monitoring infrastructure and deploy it. * Optimize system performance, cost, and security through regular reviews and tuning. ## Related Videos - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Fullstack Developer Salary UK](https://www.wearedevelopers.com/magazine/251-fullstack-developer-salary-uk)