Site Reliability Engineer
STN, inc.
United States
about 2 months ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source
Tech stack
Computer Programming
DevOps
Python (Programming Language)
Reliability Engineering
Software Reliability Testing
Prometheus
Datadog
Grafana
Mttr
Kubernetes
Job description
The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution., * Define and operate Service Level Objectives (SLOs) aligned with customer SLAs
- Build and maintain the observability stack including metrics, logs, traces, and alerting
- Lead incident response and chair post-incident reviews
- Drive automation to reduce toil and improve mean-time-to-recover (MTTR)
- Author and maintain operational runbooks alongside the NOC
- Manage on-call rotation, escalation paths, and incident-management tooling
- Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering
- Drive chaos engineering, game days, and reliability testing programs
- Produce SLA performance reports in coordination with the SLA Manager
- Mentor junior engineers and contribute to engineering culture
Requirements
Do you have experience in Tooling?, * 5+ years in SRE, DevOps, or production engineering roles
- Strong programming skills in Go, Python, or both
- Hands-on experience operating Kubernetes-based platforms at scale
- Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)
- Strong incident management experience including major-incident command, * GPU or HPC platform operational experience
- Familiarity with SLA-driven customer environments and credit calculations
- Experience with chaos engineering tools (Gremlin, Litmus, or similar)
- Published SRE content or contributions
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
about 2 years ago
EM
Eli McGarvie
Find a Developer Job: 12 Best Job Sites For Developers
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
LM
Luis Minvielle
Where To Find Software Engineering Jobs
over 2 years ago
Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!
almost 6 years ago