> Markdown version of [/jobs/ext/2067144-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2067144-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** NICE - **Location:** UK (Remote available) - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Computer Programming, Domain Name System (DNS), Python (Programming Language), Linux System Administration, Networking Basics, Reliability Engineering, Ansible, Prometheus, TCP/IP, Datadog, Scripting, Google Cloud, Load Balancing, Computer Network Operations, Grafana, Mttr, Kubernetes, Cloudwatch, Terraform, Splunk, Docker, Golang - **Published:** August 15, 2026 - **Apply:** https://www.totaljobs.com/job/site-reliability-engineer/nice-job107844089 ## About the Role * Strong Linux systems administration * Experience with incident management and production support * Familiarity with: * Cloud infrastructure (AWS preferred) * Containers & orchestration (Docker, Kubernetes) * Monitoring/alerting platforms Scripting or programming experience in Python, Bash, Go, or similar Understanding of networking fundamentals (DNS, TCP/IP, load balancing) Operational * Experience working in 24x7 NOC or production operations environments * Ability to handle high-pressure incidents calmly and effectively * Strong written and verbal communication for incident coordination * Comfort working from runbooks-but improving them when they fall short Preferred / Differentiators * Experience defining or operating to SLOs / SLIs * Prior migration from traditional NOC ? SRE model * Infrastructure as Code experience (Terraform, Ansible, etc.) * Exposure to security, compliance, or regulated environments ## Description At NiCE, we don't limit our challenges. We challenge our limits. Always. We're ambitious. We're game changers. And we play to win. We set the highest standards and execute beyond them. And if you're like us, we can offer you the ultimate career opportunity that will light a fire within you. So, what's the role all about? The SRE - NOC role sits at the intersection of traditional Network Operations Center (NOC) responsibilities and engineering-driven reliability practices. This role focuses on 24/7 service reliability, incident response, operational automation, and observability, while actively reducing operational toil through software and automation. Unlike a traditional NOC analyst, an SRE-NOC is expected to engineer problems away, not just respond to alerts. How will you make an impact? Incident Response & Operations * Act as a primary or escalation responder in a 24x7 on-call rotation * Lead or support Major Incident (MI) response, including triage, mitigation, and resolution * Coordinate across Engineering, Infrastructure, Security, and Product teams * Execute and improve runbooks, playbooks, and escalation paths * Drive blameless post-incident reviews (PIRs) and track corrective actions Monitoring, Alerting & Observability * Own service health monitoring across infrastructure, applications, and dependencies * Design and maintain alerting strategies that align with SLIs/SLOs * Reduce alert fatigue through signal-to-noise improvements * Build dashboards using tools such as: * Grafana * Prometheus * Datadog / Splunk / CloudWatch Reliability Engineering & Automation * Automate repetitive operational tasks to reduce manual toil * Improve mean time to detect (MTTD) and mean time to resolve (MTTR) * Develop scripts and tools (Python, Bash, Go, etc.) to support NOC/SRE workflows * Implement self-healing and auto-remediation where possible * Partner with engineering teams to improve system design for reliability Platform & Infrastructure Support * Support and troubleshoot: * Linux-based systems * Cloud platforms (AWS, Azure, GCP) * Kubernetes / containerized environments Assist with capacity planning and availability reviews Ensure operational readiness for production releases ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london)