> Markdown version of [/jobs/ext/611693-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/611693-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** STN, inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Computer Programming, DevOps, Python (Programming Language), Reliability Engineering, Software Reliability Testing, Prometheus, Datadog, Grafana, Mttr, Kubernetes - **Published:** June 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8f0dbe42577e0a05 ## About the Role Do you have experience in Tooling?, * 5+ years in SRE, DevOps, or production engineering roles * Strong programming skills in Go, Python, or both * Hands-on experience operating Kubernetes-based platforms at scale * Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry) * Strong incident management experience including major-incident command, * GPU or HPC platform operational experience * Familiarity with SLA-driven customer environments and credit calculations * Experience with chaos engineering tools (Gremlin, Litmus, or similar) * Published SRE content or contributions ## Description The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution., * Define and operate Service Level Objectives (SLOs) aligned with customer SLAs * Build and maintain the observability stack including metrics, logs, traces, and alerting * Lead incident response and chair post-incident reviews * Drive automation to reduce toil and improve mean-time-to-recover (MTTR) * Author and maintain operational runbooks alongside the NOC * Manage on-call rotation, escalation paths, and incident-management tooling * Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering * Drive chaos engineering, game days, and reliability testing programs * Produce SLA performance reports in coordination with the SLA Manager * Mentor junior engineers and contribute to engineering culture ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too)