> Markdown version of [/jobs/ext/1321653-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1321653-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Talkiatry Management Services, LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $160,000.0 - $185,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Python (Programming Language), Node.Js, Reliability Engineering, Prometheus, TypeScript, Datadog, Data Logging, ReactJS, Grafana, Kubernetes, Terraform, Dynatrace, AWS EKS - **Published:** July 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=2ab85dfeea6dcc0e ## About the Role * 7+ years in software or infrastructure engineering, with substantial hands-on SRE or production reliability experience. * A track record of reducing incidents and improving detection-the outcomes this role is judged on. * Hands-on experience defining SLOs/SLIs and using error budgets to guide engineering decisions. * Deep observability expertise across metrics, logging, tracing, and alerting (e.g., Datadog, Prometheus, Grafana, or similar). * Strong experience operating production systems on AWS. * Proficiency with infrastructure-as-code (e.g., Terraform) and comfort building automation and tooling (Python, TypeScript, or similar). * Excellent communication skills, with the ability to influence and align teams you don't directly manage. Nice to Have * Experience standing up an SRE function for the first time at a startup or scale-up. * Familiarity with the stack our teams run on (TypeScript/Node.js, React, AWS EKS & RDS). * Background in healthcare, tele-health, or other regulated, compliance-sensitive environments (e.g., HIPAA). * Kubernetes or container orchestration experience. ## Description This is a high-leverage, founding role. You won't inherit an SRE playbook-you'll write it. Importantly, our product teams will continue to own on-call for their own services; you're not the pager. Instead, you'll partner with engineers across patient-facing and platform teams to establish SLOs, sharpen observability, reduce toil, and shift incident detection from "a stakeholder told us" to "our monitoring caught it first." You'll act as a force multiplier, making reliability a shared responsibility rather than a separate function, and you'll measure your impact in fewer outages and calmer, quieter on-call rotations for the teams you support., * Define and roll out an SRE practice for a six-team organization: SLOs/SLIs, error budgets, and reliability standards that teams genuinely adopt. * Build and improve observability-metrics, logging, distributed tracing, dashboards, and alerting-so that more incidents are detected by monitoring before anyone outside engineering notices. * Drive down outage frequency by surfacing systemic reliability risks and partnering with teams to remediate them at the root. * Reduce toil through automation, infrastructure-as-code, and self-service tooling that teams can own and extend themselves. * Own the health and usability of our observability tooling, providing documentation and training where necessary. * Run production readiness reviews for new services and partner with engineering leadership on reliability priorities and capacity planning. ## Related Videos - [Watch Tests Go Brrrr! : Getting Started with Cypress in ReactJS](https://www.wearedevelopers.com/videos/282-watch-tests-go-brrrr-getting-started-with-cypress-in-reactjs) - [The OpenTelemetry mistakes I keep seeing (and how to stop making them)](https://www.wearedevelopers.com/videos/100158-the-opentelemetry-mistakes-i-keep-seeing-and-how-to-stop-making-them) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [React Developer Salary [2023]](https://www.wearedevelopers.com/magazine/198-react-developer-salary-2023)