> Markdown version of [/jobs/ext/622339-telecommute-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/622339-telecommute-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # TELECOMMUTE Lead Site Reliability Engineer - **Company:** Alteryx, Inc. - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $136,000.0 - $177,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), JavaScript (Programming Language), Artificial Intelligence, C++ (Programming Language), Software as a Service, Software Quality, Continuous Integration, Disaster Recovery, Distributed Systems, Failover, Python (Programming Language), Reliability Engineering, Datadog, Grafana, Mttr, Reliability of Systems, Kubernetes - **Published:** June 24, 2026 - **Apply:** https://www.dice.com/job-detail/dd63dc76-bc02-4888-a25c-bae00dd1c7ee ## About the Role 6+ years leading delivery of complex, distributed systems or SaaS platforms Strong experience with multi-region, split-plane architectures (control-plane / data-plane) Proven track record improving SLOs, MTTR, and system reliability at scale Proficiency in languages like Python, Java, C++, or JavaScript Deep experience with: Kubernetes (multi-cluster), CI/CD, and GitOps (ArgoCD) SLO/SLA design, observability, and incident management Infrastructure as Code and cloud platforms Disaster recovery, resilience, and security best practices Strong leadership skills with experience mentoring senior engineers and influencing cross-team decisions Nice to Have Experience with chaos engineering and large-scale reliability automation Background in enterprise SaaS platforms or split-plane architectures Expertise in navigating, understanding and leveraging modern Observability platfroms (Datadog, Grafana, etc) ## Description We're looking for a Lead SRE to own reliability outcomes for a modern split-plane, multi-region SaaS platform serving enterprise customers. This is a hands-on technical leadership role focused on system design, reliability strategy, and cross-team execution. You'll lead efforts that directly impact SLO attainment, MTTR reduction, and cost efficiency, while shaping how reliability is engineered, measured, and scaled across the platform. What You'll Do Define and drive reliability strategy across control-plane and data-plane systems, including multi-region resilience, BCDR, and failover design Establish and operationalize SLOs, SLAs, and error budgets, ensuring they inform planning and engineering tradeoffs Lead initiatives that measurably improve MTTR, incident prevention, and overall service health Own incident management end-to-end, driving systemic fixes and long-term reliability improvements beyond immediate response Lead architecture and design reviews to ensure systems meet scalability, reliability, and cost efficiency goals Champion automation and modernization, including AI-driven reliability improvements Establish and enforce code quality and review standards Lead cross-functional initiatives and align engineering with product priorities Mentor senior engineers and act as a technical leader across teams ## Related Videos - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Mastering Remote Work: Tips for Developers](https://www.wearedevelopers.com/magazine/558-mastering-remote-work-tips-for-developers)