> Markdown version of [/jobs/ext/2737720-sr-site-reliability-engineer-observability](https://www.wearedevelopers.com/jobs/ext/2737720-sr-site-reliability-engineer-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Site Reliability Engineer, Observability - **Company:** Tesla Motors - **Location:** Fremont, CA, United States - **Experience:** Expert - **Salary:** $120,000.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Artificial Intelligence, Amazon S3, ARM Architecture, Authentication Protocols, System Configuration, Linux, Distributed Systems, Github, Protocol Buffers, Monitoring of Systems, Python (Programming Language), OAuth, Online Transaction Processing, Performance Tuning, Reliability Engineering, Ansible, Prometheus, SQL Databases, Grafana, Reliability of Systems, Kubernetes, Deployment Automation, Apache Kafka, Api Gateway, Stream Processing, Splunk - **Published:** September 5, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/18218787?backUrl=%2Fcareer%2F18218787%2FSr-Site-Reliability-Engineer-Observability-California-Fremont ## About the Role * Strong hands-on experience with observability stacks including Grafana Mimir / Prometheus / cortex / Thanos, or equivalent enterprise-grade metrics platforms * Deep expertise in Linux system internals, large-scale performance tuning, and systems administration * Solid hands-on experience with Kubernetes configuration, networking, deployment, and multi-cluster HA architectures * Advanced proficiency in PromQL and SQL, with strong understanding of high-cardinality metrics, label design, and series explosion impacts on storage and query performance * Strong knowledge of monitoring and observability practices including OpenTelemetry (OTLP), Protobuf, and Prometheus-based metrics collection * Experience with distributed systems architecture, multi-region deployments, and high-availability cluster design * Good to have familiarity with S3-compatible object storage and exposure to distributed streaming systems such as Apache Kafka or Redpanda * Good to have knowledge of configuring and managing authentication mechanisms (OAuth, reverse proxies, API gateways, mTLS) * Proven troubleshooting expertise and performance optimization experience in large-scale distributed metrics and logs platforms. Splunk administration is a plus * Strong scripting and automation skills (Python, Ansible, GitHub Actions), excellent documentation practices, and participation in on-call and incident management processes ## Description This role requires deep expertise in observability technologies, distributed systems, Kubernetes, and high-availability design, as well as the ability to leverage AI-assisted tooling to accelerate troubleshooting, operational efficiency, and system reliability. What You'll Do * Design, build, and operate a multi-tenant Grafana Mimir estate processing 1B+ active series, including ingestion, query paths, compactors, and long-term object storage * Operate Splunk as a multi-site clustered platform ingesting 700+ TB of logs per day, covering indexer health, search head performance, and cluster replication * Build and tune Cribl (or equivalent) pipelines for routing, filtering, and enrichment of metrics and logs across sitesExtend Cardinal Wright, our internal cardinality watchdog, to detect tenant label explosions at ingestion before they impact the query path * Run Kubernetes for the observability stack itself, including node placement, PVC and storage failures, workload scheduling, and cluster monitoring * Participate in on-call rotations and lead incident response when the observability platform itself is the source of an outage * Apply AI to this stack where it is safe to do so including query helpers, anomaly detection hints, and draft runbooks while keeping a human in the loop when the blast radius could extend to a factory line or vehicle * Collaborate with SREs, architects, and application teams to deliver end-to-end service visibility across Tesla's digital, manufacturing, fleet, and Autopilot platforms ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Keeping applications secure by evolving OAuth 2.0 and OpenID Connect](https://www.wearedevelopers.com/videos/100152-keeping-applications-secure-by-evolving-oauth-2-0-and-openid-connect) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Delay the AI Overlords: How OAuth and OpenFGA Can Keep Your AI Agents from Going Rogue](https://www.wearedevelopers.com/videos/1637-delay-the-ai-overlords-how-oauth-and-openfga-can-keep-your-ai-agents-from-going-rogue) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 139 - Soft and hard queries](https://www.wearedevelopers.com/magazine/487-dev-digest-139-soft-and-hard-queries) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)