Site Reliability Engineer

STN, inc.
United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Computer Programming DevOps Python (Programming Language) Reliability Engineering Software Reliability Testing Prometheus Datadog Grafana Mttr Kubernetes

Job description

The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution., * Define and operate Service Level Objectives (SLOs) aligned with customer SLAs

  • Build and maintain the observability stack including metrics, logs, traces, and alerting
  • Lead incident response and chair post-incident reviews
  • Drive automation to reduce toil and improve mean-time-to-recover (MTTR)
  • Author and maintain operational runbooks alongside the NOC
  • Manage on-call rotation, escalation paths, and incident-management tooling
  • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering
  • Drive chaos engineering, game days, and reliability testing programs
  • Produce SLA performance reports in coordination with the SLA Manager
  • Mentor junior engineers and contribute to engineering culture

Requirements

Do you have experience in Tooling?, * 5+ years in SRE, DevOps, or production engineering roles

  • Strong programming skills in Go, Python, or both
  • Hands-on experience operating Kubernetes-based platforms at scale
  • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)
  • Strong incident management experience including major-incident command, * GPU or HPC platform operational experience
  • Familiarity with SLA-driven customer environments and credit calculations
  • Experience with chaos engineering tools (Gremlin, Litmus, or similar)
  • Published SRE content or contributions

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all