Staff Software Engineer, Site Reliability Engineer

Harvey, Inc.
San Francisco, CA, United States
11 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$238,000.0 - $290,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Cloud Computing Security Cyber Security Computer Programming Databases Continuous Delivery Continuous Integration Data Security
+17 more
Python (Programming Language) Systems Development Life Cycle Reliability Engineering Software Engineering Datadog Pulumi Scripting Google Cloud Grafana Reliability of Systems Cloudformation Containerization Kubernetes Sentry Terraform Pagerduty Golang

Job description

As a Staff Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability, and performance of our legal AI platform. You’ll join a high-leverage team that sits at the intersection of infrastructure and product, owning the systems that keep our platform fast, secure, and always on. From scaling across 50+ regions to automating mission-critical operations, your work will ensure that Harvey remains resilient as we grow. If you’re passionate about building robust systems and reducing complexity through automation, we’d love to work with you.

This role is based in San Francisco, CA. We use an in-person work model and offer relocation assistance to new employees. What You’ll Do

  • Design, implement, and manage monitoring, alerting, and infrastructure resources (compute, storage, networking) across 50+ global regions
  • Lead incident management processes, including postmortems, root cause analyses, and driving actionable improvements
  • Automate operational tasks and workflows, building tools and processes for capacity planning, graceful rollouts, and safe data access to maintain high reliability and reduce manual intervention
  • Establish best practices for security, compliance, and reliability and collaborate across teams to drive these principles throughout the software lifecycle
  • Optimize infrastructure costs through strategic capacity planning and build-versus-buy decisions while maintaining system performance, reliability, and functionality
  • Provide technical mentorship and leadership, promoting best practices and fostering team growth

Requirements

  • 10+ years of experience in Site Reliability Engineering or similar roles supporting production environments, with proven ability to mentor and guide technical teams
  • Expertise in infrastructure as code(IaC) tools (Pulumi, Terraform, CloudFormation, etc.)
  • Deep familiarity with observability tools (Datadog, Sentry, etc.) and incident response practices (PagerDuty, IncidentIO, etc.)
  • Proficiency with cloud infrastructure platforms (Azure, GCP, AWS, etc.)
  • Strong programming skills (Python, Bash, Go, or similar languages)
  • Proven track record of diagnosing complex system problems and implementing durable solutions
  • Solid understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles
  • Excellent problem-solving skills, meticulous attention to detail, and a commitment to operational excellence, Artificial Intelligence (AI), Automation, Bash Scripting, Best Practices, Capacity Management, Category Development, Cloud Computing, Computer Programming, Continuous Deployment/Delivery, Continuous Integration, Cost Control, Detail Oriented, Financial Services, Go Programming Language (Golang), High Reliability, Identify Issues, Incident Management, Incident Response, Information/Data Security (InfoSec), Leadership, Legal, Mentoring, Problem Solving Skills, Process Improvement, Production Support, Production Systems, Professional Services, Python Programming/Scripting Language, Reliability Engineering, Root Cause Analysis, Software Development Lifecycle (SDLC), Software Engineering, Strategic Planning, Systems Maintenance, Systems Reliability

About the company

Title: Staff Software Engineer, Site Reliability Engineer Organization: Harvey Location: San Francisco Description: Why Harvey

At Harvey, we’re transforming how legal and professional services operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for decades to come.

This is a rare chance to help build a generational company at a true inflection point. With 2400+ customers in 70+ countries, strong product-market fit, and world-class investor support, we’re scaling fast and defining a new category in real time. The work is ambitious, the bar is high, and the opportunity for growth - personal, professional, and financial - is unmatched.

Our team moves fast, takes ownership, and is deeply committed to the mission - operating with intensity, staying close to our customers, and pushing each other for excellence. We live by three values: Decisiveness, Simplicity, and Job’s Not Finished. We act quickly on clear judgment over perfect information, we believe simplicity is what scales, and we’re never satisfied with where we are. If you want to do the best work of your career alongside people who share that drive, we’d love to build with you.

At Harvey, the future of professional services is being written today - and we’re just getting started. Role Overview

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

1:22 min

Overview of the Sentry error and performance monitoring platform

Priscila Oliveira · WWC 2023

1:33 min

Case study on adopting Kubernetes and Golang effectively

Andrew Holway · LIVE

3:20 min

Overview of infrastructure as code tools

Alexander Bubeck · WWC 2023

Videos

See all

Related articles

See all