Site Reliability Engineer

Obsidian Security
Salford, UK
7 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Software as a Service Cloud Computing DevOps Distributed Systems Reliability Engineering Scripting Delivery Pipeline Kubernetes Infrastructure Automation Frameworks Virtual Agents Data Pipelines
+2 more
Programming Languages Microservices

Job description

At Obsidian, our Site Reliability Engineers ensure the reliability, scalability, and operational excellence of a complex multi-tenant SaaS platform serving enterprise and financial customers. As an SRE, you will work closely with DevOps, Platform Engineering, and product teams to improve system observability, incident response, and service resilience across the platform., * Reliability Engineering: Improve the reliability, availability, and resiliency of Obsidian’s production systems and distributed services

  • Detection & Observability: Build and maintain monitoring, alerting, dashboards, and observability tooling to enhance system visibility and reduce operational noise
  • Incident Response & Operations: Support incident response, on-call operations, troubleshooting, and postmortem processes to drive operational excellence
  • Collaboration: Partner with engineering teams to implement SLI/SLO practices, operational standards, and reliability-focused workflows
  • Execution: Automate infrastructure operations, deployment workflows, and platform tooling across Kubernetes, cloud infrastructure, and data pipelines, * Work on reliability challenges across a large-scale distributed SaaS platform
  • Build and improve observability and operational tooling used across engineering
  • Gain hands-on experience with cloud infrastructure, Kubernetes, and production systems
  • Help safeguard critical services for enterprise and financial customers

What Success Looks Like

  • Production issues are detected and resolved quickly
  • Monitoring and alerting provide clear, actionable operational insights
  • Reliability metrics and operational practices improve over time
  • Engineering teams can effectively troubleshoot and self-serve observability
  • Automation reduces operational toil and improves platform stability

Requirements

  • 2-5 years of experience in Site Reliability Engineering, DevOps, Production Engineering, or related roles
  • Experience operating and supporting production systems in AWS and/or GCP
  • Familiarity with Kubernetes and Helm in cloud-native environments
  • Experience with observability and monitoring tools such as Prometheus, Grafana, Datadog, or similar platforms
  • Exposure to CI/CD systems such as GitLab CI/CD, GitHub Actions, ArgoCD, or equivalent
  • Strong troubleshooting and debugging skills across distributed systems and microservices
  • Experience writing automation or infrastructure tooling using scripting or programming languages
  • Strong systems thinking and a collaborative engineering mindset, * AI Agent development experience
  • Experience supporting SaaS platforms in production environments
  • Familiarity with incident management and postmortem practices
  • Exposure to infrastructure-as-code and GitOps workflows
  • Understanding of SLI/SLO concepts and operational metrics
  • Experience with enterprise-scale monitoring or customer-facing production systems

Benefits & conditions

  • Competitive compensation with equity and 401k
  • Comprehensive healthcare with dental and vision coverage
  • Flexible paid time off and paid holiday time off
  • 12 weeks of new parent or family leave
  • Personal and professional development resources

For more details on our US benefits, or for information on our international benefits, please see here.

Pay Transparancy Please note that the base pay range is a guideline and for candidates who receive an offer, the base pay will vary based on factors such as work location, as well as the knowledge, skills and experience of the candidate. In addition to a competitive base salary, this position is eligible for equity awards and may be eligible for sales commission or incentive compensation based on the role or function within the company.

At Obsidian, we are proud to be an equal-opportunity employer. We value diversity and hire for talent, passion, and compassion. In compliance with federal law, all persons hired will be required to submit satisfactory proof of identity and legal authorization. If you have a need that requires accommodation, please contact

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all