Staff Software Engineer, Infrastructure Engineering

Tesla Motors
Austin, TX, United States
5 days ago
Apply on jobs.localjobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Common ISDN Application Programming Interface (CAPI) Programming Tools Distributed Computing Environment Distributed Systems Job Scheduling Python (Programming Language) Network Control Azure Machine Learning Service-Oriented Architecture Cloud Platform System
+8 more
Pytorch Autoscaling Backend Kubernetes Machine Learning Operations Hardware Infrastructure Software Version Control Serverless Computing

Job description

You will be responsible for the internal platform that lets every team securely spin up sandboxed AI agents, run ephemeral workloads, train models at scale, and serve them in production all as a reliable, self-service experience. This is a high-leverage, high-ownership role where you will design, build, and operate the systems that sit at the intersection of Kubernetes, high-performance networking, GPU infrastructure, and modern ML platforms. What You’ll Do

  • Build and own the end-to-end AI/ML platform (training, inference, experimentation) as a self-service product for all internal users
  • Write production Kubernetes operators and controllers in Go for GPU workloads, training jobs, model deployments, sandboxes, and cluster lifecycle
  • Operate large-scale GPU fleets (A100/H100/B200) scheduling, MIG, topology-aware placement, health monitoring
  • Build and operate inference infrastructure using KServe, Triton, vLLM, and Ray Serve autoscaling, model versioning, and request batching
  • Build and operate training-as-a-service; distributed training (PyTorch , FSDP), MLflow, and checkpoint management
  • Design secure, isolated sandbox environments for AI agents and ephemeral execution contexts for untrusted code (gVisor, Kata, Firecracker)
  • Architect serverless/Lambda-style ephemeral workloads, scale-to-zero, event-driven compute (Knative, KEDA, Firecracker microVMs)
  • Integrate GPU compute, training, inference, and sandboxes as first-class services in the internal cloud platform
  • Manage Kubernetes cluster lifecycle at fleet scale using Cluster API (CAPI) provisioning, upgrades, scaling, and decommissioning

Requirements

  • Have a minimum of 8+ years of practical experience as a backend or platform engineer building scalable distributed systems, ideally operating at tech lead level , cloud-native products, developer tools, or external developer facing products
  • Built agent sandbox/ephemeral compute platforms (gVisor, Kata, Firecracker, or similar)
  • Strong fundamentals in service-oriented architectures, networking, and systems design, ideally having owned both the technical vision and execution of a foundational platform system end to end
  • Strong profciency in Python, Go, Rust, or similar systems languages (full stack)
  • Production experience with KServe, Triton, vLLM, or Ray Serve for model inference
  • Built or operated training-as-a-service, distributed training, MLflow, job scheduling
  • Root cause analysis and systems thinking - failure modes, back-pressure, resource contention, blast radius
  • Deep Kubernetes internals expertise (scheduler, API server, etcd, admission controllers, CRDs, control plane at scale)
  • Built production Kubernetes operators in Go (controller-runtime / Kubebuilder)
  • Production Cluster API (CAPI) experience; management clusters, custom providers, ClusterClass, fleet-scale lifecycle

Benefits & conditions

Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire:

  • Medical plans > plan options with $0 payroll deduction
  • Family-building, fertility, adoption and surrogacy benefits
  • Dental (including orthodontic coverage) and vision plans, both have options with a $0 paycheck contribution
  • Company Paid (Health Savings Accounts) HSA Contribution when enrolled in the High-Deductible medical plan with HSA
  • Healthcare and Dependent Care Flexible Spending Accounts (FSA)
  • 401(k) with employer match, Employee Stock Purchase Plans, and other financial benefits
  • Company paid Basic Life, AD&D
  • Short-term and long-term disability insurance (90 day waiting period)
  • Employee Assistance Program
  • Sick and Vacation time (Flex time for salary positions, Accrued hours for Hourly positions), and Paid Holidays
  • Back-up childcare and parenting support resources
  • Voluntary benefits to include: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
  • Weight Loss and Tobacco Cessation Programs
  • Tesla Babies program
  • Commuter benefits
  • Employee discounts and perks program

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.localjobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:38 min

Transitioning into backend engineering from web development

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all