> Markdown version of [/jobs/ext/2307981-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2307981-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Obsidian Security - **Location:** Salford, UK - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Software as a Service, Cloud Computing, DevOps, Distributed Systems, Reliability Engineering, Scripting, Delivery Pipeline, Kubernetes, Infrastructure Automation Frameworks, Virtual Agents, Data Pipelines, Programming Languages, Microservices - **Published:** August 30, 2026 - **Apply:** https://www.collegerecruiter.com/job/2815118131-site-reliability-engineer ## About the Role * 2-5 years of experience in Site Reliability Engineering, DevOps, Production Engineering, or related roles * Experience operating and supporting production systems in AWS and/or GCP * Familiarity with Kubernetes and Helm in cloud-native environments * Experience with observability and monitoring tools such as Prometheus, Grafana, Datadog, or similar platforms * Exposure to CI/CD systems such as GitLab CI/CD, GitHub Actions, ArgoCD, or equivalent * Strong troubleshooting and debugging skills across distributed systems and microservices * Experience writing automation or infrastructure tooling using scripting or programming languages * Strong systems thinking and a collaborative engineering mindset, * AI Agent development experience * Experience supporting SaaS platforms in production environments * Familiarity with incident management and postmortem practices * Exposure to infrastructure-as-code and GitOps workflows * Understanding of SLI/SLO concepts and operational metrics * Experience with enterprise-scale monitoring or customer-facing production systems ## Description At Obsidian, our Site Reliability Engineers ensure the reliability, scalability, and operational excellence of a complex multi-tenant SaaS platform serving enterprise and financial customers. As an SRE, you will work closely with DevOps, Platform Engineering, and product teams to improve system observability, incident response, and service resilience across the platform., * Reliability Engineering: Improve the reliability, availability, and resiliency of Obsidian's production systems and distributed services * Detection & Observability: Build and maintain monitoring, alerting, dashboards, and observability tooling to enhance system visibility and reduce operational noise * Incident Response & Operations: Support incident response, on-call operations, troubleshooting, and postmortem processes to drive operational excellence * Collaboration: Partner with engineering teams to implement SLI/SLO practices, operational standards, and reliability-focused workflows * Execution: Automate infrastructure operations, deployment workflows, and platform tooling across Kubernetes, cloud infrastructure, and data pipelines, * Work on reliability challenges across a large-scale distributed SaaS platform * Build and improve observability and operational tooling used across engineering * Gain hands-on experience with cloud infrastructure, Kubernetes, and production systems * Help safeguard critical services for enterprise and financial customers What Success Looks Like * Production issues are detected and resolved quickly * Monitoring and alerting provide clear, actionable operational insights * Reliability metrics and operational practices improve over time * Engineering teams can effectively troubleshoot and self-serve observability * Automation reduces operational toil and improves platform stability ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Intermediate Bitcoin Script](https://www.wearedevelopers.com/videos/25-intermediate-bitcoin-script) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)