Staff Site Reliability Engineer

Filevine, Inc.
United States
8 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$235,000.0 - $275,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Bash Shell Distributed Systems Python (Programming Language) Reliability Engineering Software Engineering Datadog Kubernetes New Relic (SaaS) Golang

Job description

As a Staff Site Reliability Engineer at Filevine, you are the senior technical authority on the SRE team and a strategic partner to engineering leadership. You don’t just maintain systems - you shape engineering culture, define the technical standard for how Filevine runs in production, and bridge the gap between high-level business goals and robust, internet-scale technical execution. You bring a forward-looking perspective - actively shaping how AI and machine learning drive the future of reliability practice. You own the roadmap across two critical SRE domains - Observability & Alerting and Platform Infrastructure - and are accountable for ensuring the team solves reliability problems permanently rather than absorbing them as toil. You operate as the senior IC counterpart to the Engineering Manager: technical correctness lives with you. You partner with the Reliability Architect and engineering leadership on significant technical decisions, mentor engineers across experience levels, and influence reliability strategy across the broader organization. Reliability at Filevine protects revenue. You are the senior technical voice responsible for ensuring that uptime, incident response, and every production change meet the operational standard the business demands. This role does not participate in on-call rotation, but you are deeply invested in the engineers who do - shaping the on-call strategy, tooling, and culture that make production support sustainable and effective., Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.

  • Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed

systems.

  • Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation

across the service lifecycle.

  • Lead the organization through complex production incidents and turn post-incident learning into

permanent engineering improvements.

  • Build self-service platform capabilities that reduce toil, improve engineering safety and velocity,

and make every team more capable of owning their own reliability.

  • Mentor engineers and serve as a trusted technical authority for long-term reliability and platform

direction.

Requirements

observability, and reliability engineering. You raise the technical standard for every engineer around you and thrive where the challenges are complex and the stakes are real.

  • Technical Leader and Mentor: You are passionate about mentoring engineers and investing in

their growth. You influence technical direction and communicate production risk clearly across engineering, product, and executive audiences.

  • Forward-Thinking & AI/ML Fluent: You bring deep knowledge of AIOps and drive the use of

AI and machine learning in observability, anomaly detection, incident response, automated remediation, and resource optimization.

  • Production-Scale Problem Solver: You turn ambiguous, complex reliability challenges into

durable solutions for systems where availability, performance, and production changes carry meaningful business impact.

  • Software-Minded Builder: You use software, automation, Infrastructure as Code, and platform, 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE, including 6+ years in SRE and 3+ years leading complex, cross-functional technical initiatives for distributed production systems.

  • Expert-level depth in observability and platform infrastructure, with broad expertise in incident

response, capacity planning, automation, and reliability engineering.

  • Advanced experience with a major container-orchestration platform, preferably Kubernetes, and

an observability platform such as New Relic, Datadog, or equivalent.

  • Strong software-engineering ability in Python, Go, Bash, or another general-purpose language,

with experience building production tooling, automation, or platform capabilities.

  • Proven ability to mentor engineers and communicate technical risk clearly to engineering,

product, and executive audiences.

  • Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is

Benefits & conditions

2.42.4 out of 5 stars Remote $235,000 - $275,000 a year - Full-time, Pulled from the full job description

  • Parental leave
  • Vision insurance
  • Dental insurance
  • Disability insurance, strongly preferred. Cool Company Benefits:

  • A dynamic, rapidly growing company, focused on helping organizations thrive

  • Medical, Dental, & Vision Insurance (for full-time employees)

  • Competitive & Fair Pay

  • Maternity & paternity leave (for full-time employees)

About the company

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. Grounded in a singular system of truth, Filevine brings together data, documents, workflows, and teams into one unified platform-where modern legal work happens with clarity and consistency. Powered by LOIS, the Legal Operating Intelligence System, Filevine connects context across every matter to transform legal operations from reactive to proactive. LOIS reads, understands, and reasons across your data to surface insight, automate complexity, and give professionals the clarity and confidence to see more, know more, and do more. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:52 min

Avoiding remote code execution from unsanitized inputs

Alexander Pirker · WWC 2022

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

54 sec

Interpreting complex terminal commands safely using external explanation utilities

Dan Cranney +2 · LIVE

Videos

See all

Related articles

See all