Sr. Site Reliability Engineer, Observability

Tesla Motors
Fremont, CA, United States
4 days ago
Apply on diversityjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$120,000.0
Working hours
Regular working hours

Tech stack

Query Performance Artificial Intelligence Amazon S3 ARM Architecture Authentication Protocols System Configuration Linux Distributed Systems Github Protocol Buffers Monitoring of Systems Python (Programming Language)
+15 more
OAuth Online Transaction Processing Performance Tuning Reliability Engineering Ansible Prometheus SQL Databases Grafana Reliability of Systems Kubernetes Deployment Automation Apache Kafka Api Gateway Stream Processing Splunk

Job description

This role requires deep expertise in observability technologies, distributed systems, Kubernetes, and high-availability design, as well as the ability to leverage AI-assisted tooling to accelerate troubleshooting, operational efficiency, and system reliability. What You’ll Do

  • Design, build, and operate a multi-tenant Grafana Mimir estate processing 1B+ active series, including ingestion, query paths, compactors, and long-term object storage
  • Operate Splunk as a multi-site clustered platform ingesting 700+ TB of logs per day, covering indexer health, search head performance, and cluster replication
  • Build and tune Cribl (or equivalent) pipelines for routing, filtering, and enrichment of metrics and logs across sitesExtend Cardinal Wright, our internal cardinality watchdog, to detect tenant label explosions at ingestion before they impact the query path
  • Run Kubernetes for the observability stack itself, including node placement, PVC and storage failures, workload scheduling, and cluster monitoring
  • Participate in on-call rotations and lead incident response when the observability platform itself is the source of an outage
  • Apply AI to this stack where it is safe to do so including query helpers, anomaly detection hints, and draft runbooks while keeping a human in the loop when the blast radius could extend to a factory line or vehicle
  • Collaborate with SREs, architects, and application teams to deliver end-to-end service visibility across Tesla’s digital, manufacturing, fleet, and Autopilot platforms

Requirements

  • Strong hands-on experience with observability stacks including Grafana Mimir / Prometheus / cortex / Thanos, or equivalent enterprise-grade metrics platforms
  • Deep expertise in Linux system internals, large-scale performance tuning, and systems administration
  • Solid hands-on experience with Kubernetes configuration, networking, deployment, and multi-cluster HA architectures
  • Advanced proficiency in PromQL and SQL, with strong understanding of high-cardinality metrics, label design, and series explosion impacts on storage and query performance
  • Strong knowledge of monitoring and observability practices including OpenTelemetry (OTLP), Protobuf, and Prometheus-based metrics collection
  • Experience with distributed systems architecture, multi-region deployments, and high-availability cluster design
  • Good to have familiarity with S3-compatible object storage and exposure to distributed streaming systems such as Apache Kafka or Redpanda
  • Good to have knowledge of configuring and managing authentication mechanisms (OAuth, reverse proxies, API gateways, mTLS)
  • Proven troubleshooting expertise and performance optimization experience in large-scale distributed metrics and logs platforms. Splunk administration is a plus
  • Strong scripting and automation skills (Python, Ansible, GitHub Actions), excellent documentation practices, and participation in on-call and incident management processes

Benefits & conditions

paid holidays, flex time, 401(k) United States, California, Fremont Sep 05, 2026 What to Expect In this role, you will design and operate Tesla’s enterprise-grade observability solutions, delivering end-to-end visibility across digital, manufacturing, fleet, and Autopilot platforms. You will be part of the Observability team, which owns and operates the full observability stack, including a large-scale metrics platform processing more than one billion active time series, dashboards and alerting for real-time visibility, and a log ingestion platform handling 700+ TB of data daily. Together, these systems provide critical visibility across Tesla’s internal and global infrastructure., Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire:

  • Medical plans > plan options with $0 payroll deduction
  • Family-building, fertility, adoption and surrogacy benefits
  • Dental (including orthodontic coverage) and vision plans, both have options with a $0 paycheck contribution
  • Company Paid (Health Savings Accounts) HSA Contribution when enrolled in the High-Deductible medical plan with HSA
  • Healthcare and Dependent Care Flexible Spending Accounts (FSA)
  • 401(k) with employer match, Employee Stock Purchase Plans, and other financial benefits
  • Company paid Basic Life, AD&D
  • Short-term and long-term disability insurance (90 day waiting period)
  • Employee Assistance Program
  • Sick and Vacation time (Flex time for salary positions, Accrued hours for Hourly positions), and Paid Holidays
  • Back-up childcare and parenting support resources
  • Voluntary benefits to include: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
  • Weight Loss and Tobacco Cessation Programs
  • Tesla Babies program
  • Commuter benefits
  • Employee discounts and perks program

Expected Compensation $120,000 - $396,000/annual salary + cash and stock awards + benefits

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:49 min

Adopting OAuth best practices and removing outdated grants

Alexander Schwartz Alexander Schwartz · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

1:34 min

Analyzing vulnerabilities in standard OAuth 2.0 authorization flows

Alexander Schwartz Alexander Schwartz · World Congress 2026 Europe

Videos

See all

Related articles

See all