> Markdown version of [/jobs/ext/2806456-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2806456-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Incite Insight - **Location:** London, UK - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, JIRA, Data Centers, Data Center Infrastructure Management (CIM), DevOps, Monitoring of Systems, Python (Programming Language), Reliability Engineering, Prometheus, Software Engineering, Systems Integration, Large Language Models, Grafana, Slack, Hardware Infrastructure, Dynatrace, Servicenow - **Published:** September 9, 2026 - **Apply:** https://www.collegerecruiter.com/job/2858056400-site-reliability-engineer ## About the Role utilities, dashboards and potentially ChatOps capabilities that allow operational teams to resolve issues more quickly.A major part of the role will be taking existing operational processes and asking:\"Why are we still doing this manually?\"You'll then design a safe, controlled and auditable way of automating it.What we're looking for You should have good commercial experience in Site Reliability Engineering, Platform Engineering, DevOps or production infrastructure operations, together with strong hands-on Python automation skills.You'll also need experience with: Monitoring and observability tools such as Prometheus, Grafana or similar Production incident management and/or on-call environments Automating operational runbooks and repetitive infrastructure processes APIs and systems integration Version-controlled automation and operational tooling Experience with any of the following would be particularly useful:ServiceNow, Halo, Jira Service Management, OpenTelemetry, distributed tracing, Slack/Teams automation, datacentre or colocation environments, GPU infrastructure, DCIM, IPAM, virtualisation platforms or LLM-assisted operational automation.This is not an AI/ML development position. We're looking for someone who understands production infrastructure and can use software engineering and automation to make that infrastructure more reliable.You'll be joining a growing organisation where you'll have considerable autonomy and the opportunity to help shape the SRE and operational automation capability rather than simply inherit an established environment.Salary: TBCLocation / hybrid working: TBC #J-18808-Ljbffr ## Description Site Reliability / Platform Engineer We are recruiting for a growing technology infrastructure business that is building a new operational capability to support large-scale, high-performance computing environments.This is an excellent opportunity for an experienced Site Reliability Engineer or Platform Engineer who enjoys automating things rather than repeatedly fixing them manually.The role sits at the intersection of infrastructure, operations and software engineering. You will use Python and modern automation techniques to improve reliability, reduce manual workload and make incident response faster and more effective.What you'll be doing You'll build Python-based automation around incident management, operational runbooks and routine infrastructure tasks.You'll integrate monitoring, infrastructure and ITSM platforms through APIs, helping improve the quality of alerts through better correlation, enrichment, suppression and deduplication.You'll also develop internal tools, command-line ## Related Videos - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Stack Overflow: Community and AI](https://www.wearedevelopers.com/videos/600-stack-overflow-community-and-ai) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers)