TELECOMMUTE Lead Site Reliability Engineer

Alteryx, Inc.
San Francisco, CA, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$136,000.0 - $177,000.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) JavaScript (Programming Language) Artificial Intelligence C++ (Programming Language) Software as a Service Software Quality Continuous Integration Disaster Recovery Distributed Systems Failover Python (Programming Language) Reliability Engineering
+5 more
Datadog Grafana Mttr Reliability of Systems Kubernetes

Job description

We’re looking for a Lead SRE to own reliability outcomes for a modern split-plane, multi-region SaaS platform serving enterprise customers. This is a hands-on technical leadership role focused on system design, reliability strategy, and cross-team execution.

You’ll lead efforts that directly impact SLO attainment, MTTR reduction, and cost efficiency, while shaping how reliability is engineered, measured, and scaled across the platform.

What You’ll Do

Define and drive reliability strategy across control-plane and data-plane systems, including multi-region resilience, BCDR, and failover design Establish and operationalize SLOs, SLAs, and error budgets, ensuring they inform planning and engineering tradeoffs Lead initiatives that measurably improve MTTR, incident prevention, and overall service health Own incident management end-to-end, driving systemic fixes and long-term reliability improvements beyond immediate response Lead architecture and design reviews to ensure systems meet scalability, reliability, and cost efficiency goals Champion automation and modernization, including AI-driven reliability improvements Establish and enforce code quality and review standards Lead cross-functional initiatives and align engineering with product priorities Mentor senior engineers and act as a technical leader across teams

Requirements

6+ years leading delivery of complex, distributed systems or SaaS platforms Strong experience with multi-region, split-plane architectures (control-plane / data-plane) Proven track record improving SLOs, MTTR, and system reliability at scale Proficiency in languages like Python, Java, C++, or JavaScript Deep experience with: Kubernetes (multi-cluster), CI/CD, and GitOps (ArgoCD) SLO/SLA design, observability, and incident management Infrastructure as Code and cloud platforms Disaster recovery, resilience, and security best practices Strong leadership skills with experience mentoring senior engineers and influencing cross-team decisions

Nice to Have Experience with chaos engineering and large-scale reliability automation Background in enterprise SaaS platforms or split-plane architectures Expertise in navigating, understanding and leveraging modern Observability platfroms (Datadog, Grafana, etc)

Benefits & conditions

Alteryx is committed to fair, equitable, and transparent compensation. Final compensation will be determined by various factors such as your relevant work experience, education, certifications, skills, and geographic location.

The salary range for this role in the United States is $136,000 - $177,000.

Employees may also be eligible for a wide range of other benefits, such as a bonus or commission, medical, retirement, financial, wellness, time off, employee discounts, and others.

Interested? Learn more and apply today at alteryx.com/careers!

Find yourself checking a lot of these boxes but doubting whether you should apply? At Alteryx, we support a growth mindset for our associates through all stages of their careers. If you meet some of the requirements and you share our values, we encourage you to apply. As part of our ongoing commitment to a diverse, equitable, and inclusive workplace, we’re invested in building teams with a wide variety of backgrounds, identities, and experiences.

Benefits & Perks:

Alteryx has amazing benefits for all Associates which can be viewed here.

For roles in San Francisco and Los Angeles: Pursuant to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Alteryx will consider for employment qualified applicants with arrest and conviction records.

This position involves access to software/technology that is subject to U.S. export controls. Any job offer made will be contingent upon the applicant’s capacity to serve in compliance with U.S. export controls.

About the company

We’re living through a once-in-a-generation shift in how work gets done. Data, automation, and AI are quickly becoming the center of every business decision - and Alteryx is leading the transformation.

You’ll be working on the challenges that sit at the heart of modern business. No matter your role, the work you do will help organizations move faster, see more clearly, and tackle questions that used to feel impossible.

If you’re ready to meet the moment with innovation, curiosity, and excellence, there’s a place for you here.

Why work for just any analytics company? At Alteryx, Inc., we are explorers, dreamers and innovators. We’re on a journey to build the best analytics platform in the world, but we can’t do it without people like you, leading the way. Forget the stereotypical tech companies of the past. Embrace the unconventional, exercise your imagination and help alter the future with Alteryx.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · WWC 2025

Videos

See all

Related articles

See all