Site Reliability Engineer

Source Inc.
United States
8 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Java (Programming Language) .NET Framework PHP (Programming Language) Microsoft Azure Software as a Service Cloud Computing Reliability Engineering Prometheus Software Engineering Datadog Cloud Platform System Grafana
+2 more
Kubernetes Infrastructure Automation Frameworks

Job description

The Senior Site Reliability Engineer is responsible for the reliability, availability, scalability, and operational excellence of our SaaS platform. This role combines software engineering, cloud infrastructure, automation, and operational leadership to build resilient systems that enable rapid product delivery., * Own the reliability and performance of production services.

  • Define and operate SLIs, SLOs, and error budgets with engineering teams.
  • Lead incident response and drive blameless postmortems and continuous improvement.
  • Automate operational processes and reduce manual toil through engineering.
  • Build and operate cloud-native platforms using Azure, Kubernetes, and Infrastructure as Code.
  • Develop observability through effective monitoring, alerting, and telemetry.
  • Mentor engineers and promote reliability best practices across the organisation.

Requirements

  • Proven experience as a Senior or experienced Site Reliability Engineer with a software engineering background.
  • Ability to diagnose and make safe changes to PHP and Java or .NET applications.
  • Experience operating large-scale production SaaS systems.
  • Strong knowledge of SRE principles, incident management, observability, and operational excellence.
  • Hands-on experience with Azure, Kubernetes, Infrastructure as Code, and monitoring platforms such as Prometheus, Grafana, or Datadog.
  • Experience influencing engineering teams and driving reliability improvements through collaboration and technical leadership.

Key Attributes

  • Passion for building reliable production systems.
  • Software engineering mindset with a focus on automation.
  • Strong technical judgement and ownership.
  • Excellent communication during incidents and day-to-day collaboration.
  • Commitment to blameless culture and continuous improvement.

By applying for this role, you consent to SGI contacting you by telephone regarding recruitment services, market updates, and relevant business opportunities

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on computerjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · WWC 2025

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov · LIVE

4:47 min

Automating frontend performance metrics with Google Lighthouse

Miki Lombardi · JS Congress

Videos

See all

Related articles

See all