Sr. Site Reliability Engineer, Robotaxi Service

Tesla Motors
Austin, TX, United States
11 days ago
Apply on jobs.localjobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Artificial Intelligence Software Debugging Linux Distributed Systems Domain Name System (DNS) Log Analysis Networking Basics Octopus Deploy Reliability Engineering Ansible TCP/IP Virtual Local Area Networks
+5 more
Load Balancing Bug Reporting Information Technology Deployment Automation Terraform

Job description

This role sits within IT Infrastructure, embedded in the Robotaxi operational ecosystem across the US. What You’ll Do

  • Build relationships across Autopilot, fleet networking, platform engineering, and service engineering - becoming the person who understands how the pieces fit and who owns what
  • Identify where the critical path to reliability and scale is blocked - observability gaps, fragile integrations, processes that break under load - and drive the fix. Translate operational symptoms from the Robotaxi Ops team into well-scoped engineering problems and bring the right teams to a shared solution
  • Take on the hardest cross-system issues - the ones that don’t fit neatly into one team’s scope - working through them directly with Autopilot, fleet networking, platform, and infrastructure teams
  • Use AI-assisted tools for log analysis, anomaly detection, and root cause correlation to reduce time-to-resolution on complex, multi-system issues
  • Escalate well-scoped problems with clear reproduction steps, supporting data, and impact assessment
  • Build monitoring, alerting, and dashboards for Robotaxi-critical services - fleet health, tele-operation endpoints, service routing pipelines - integrated with on-call systems
  • Write automation that reduces toil: fleet health checks, service routing validation, and operational state monitoring
  • Write runbooks that give Ops and on-call engineers clear, actionable steps when things go wrong
  • Define priority levels for Robotaxi - full fleet connectivity, tele-operations, data uploads and OTAs. Build the severity classification framework before the incidents arrive
  • Work closely with Tesla’s Incident Management team as the Robotaxi subject matter expert - ensuring alerts route correctly, on-call coverage is scoped appropriately, and postmortem findings feed back into engineering

Requirements

  • 3+ yearsof SRE, infrastructure engineering, or production operations in a distributed systems environment
  • Solid Linux and networking fundamentals (TCP/IP, DNS, load balancing, VLAN/MTU)
  • Hands-onKubernetesin production - pod scheduling, ingress controllers, cluster diagnostics
  • Observability toolingexperience - log aggregation, metrics, or APM. You build alerts that fire when they should
  • Proficiency inPython or Gofor scripting, automation, and tooling
  • Comfort withAI-assisted toolsfor debugging, log analysis, and documentation
  • Cellular, Starlink, or Wi-Fibehavior at scale in mobile or field-deployed environments
  • Infrastructure-as-code and GitOps tooling (Ansible, Terraform, ArgoCD, Helm)
  • Background inIoT, connected vehicle, or fleet management- telemetry pipelines, OTA, vehicle-to-cloud

Benefits & conditions

Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire:

  • Medical plans > plan options with $0 payroll deduction
  • Family-building, fertility, adoption and surrogacy benefits
  • Dental (including orthodontic coverage) and vision plans, both have options with a $0 paycheck contribution
  • Company Paid (Health Savings Accounts) HSA Contribution when enrolled in the High-Deductible medical plan with HSA
  • Healthcare and Dependent Care Flexible Spending Accounts (FSA)
  • 401(k) with employer match, Employee Stock Purchase Plans, and other financial benefits
  • Company paid Basic Life, AD&D
  • Short-term and long-term disability insurance (90 day waiting period)
  • Employee Assistance Program
  • Sick and Vacation time (Flex time for salary positions, Accrued hours for Hourly positions), and Paid Holidays
  • Back-up childcare and parenting support resources
  • Voluntary benefits to include: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
  • Weight Loss and Tobacco Cessation Programs
  • Tesla Babies program
  • Commuter benefits
  • Employee discounts and perks program

About the company

Tesla’s Robotaxi service is one of the most operationally complex, safety-critical products we’ve ever built. The Cybercab fleet runs 24/7 across the US, backed by tightly coupled systems spanning vehicle connectivity, cloud backends, network infrastructure, autonomy software, and service operations.

We’re looking for an SREwho operates like a task force engineer - self-organizing around the hardest problems, finding the critical path to reliability and scale, and driving solutions across team boundaries without waiting to be assigned. When something breaks, you dig in, identify which layer is at fault, and engineer your way to a fix. You’ll be writing code, building tooling, and driving systemic improvements across a production system with no human driver as a fallback.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.localjobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:27 min

Structuring agile and interdisciplinary engineering teams

Oliver Zimmert · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

Videos

See all

Related articles

See all