> Markdown version of [/jobs/ext/2296265-staff-software-engineer-infrastructure-engineering](https://www.wearedevelopers.com/jobs/ext/2296265-staff-software-engineer-infrastructure-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, Infrastructure Engineering - **Company:** Tesla Motors - **Location:** Austin, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Common ISDN Application Programming Interface (CAPI), Programming Tools, Distributed Computing Environment, Distributed Systems, Job Scheduling, Python (Programming Language), Network Control, Azure Machine Learning, Service-Oriented Architecture, Cloud Platform System, Pytorch, Autoscaling, Backend, Kubernetes, Machine Learning Operations, Hardware Infrastructure, Software Version Control, Serverless Computing - **Published:** August 29, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88136209/1 ## About the Role * Have a minimum of 8+ years of practical experience as a backend or platform engineer building scalable distributed systems, ideally operating at tech lead level , cloud-native products, developer tools, or external developer facing products * Built agent sandbox/ephemeral compute platforms (gVisor, Kata, Firecracker, or similar) * Strong fundamentals in service-oriented architectures, networking, and systems design, ideally having owned both the technical vision and execution of a foundational platform system end to end * Strong profciency in Python, Go, Rust, or similar systems languages (full stack) * Production experience with KServe, Triton, vLLM, or Ray Serve for model inference * Built or operated training-as-a-service, distributed training, MLflow, job scheduling * Root cause analysis and systems thinking - failure modes, back-pressure, resource contention, blast radius * Deep Kubernetes internals expertise (scheduler, API server, etcd, admission controllers, CRDs, control plane at scale) * Built production Kubernetes operators in Go (controller-runtime / Kubebuilder) * Production Cluster API (CAPI) experience; management clusters, custom providers, ClusterClass, fleet-scale lifecycle ## Description You will be responsible for the internal platform that lets every team securely spin up sandboxed AI agents, run ephemeral workloads, train models at scale, and serve them in production all as a reliable, self-service experience. This is a high-leverage, high-ownership role where you will design, build, and operate the systems that sit at the intersection of Kubernetes, high-performance networking, GPU infrastructure, and modern ML platforms. What You'll Do * Build and own the end-to-end AI/ML platform (training, inference, experimentation) as a self-service product for all internal users * Write production Kubernetes operators and controllers in Go for GPU workloads, training jobs, model deployments, sandboxes, and cluster lifecycle * Operate large-scale GPU fleets (A100/H100/B200) scheduling, MIG, topology-aware placement, health monitoring * Build and operate inference infrastructure using KServe, Triton, vLLM, and Ray Serve autoscaling, model versioning, and request batching * Build and operate training-as-a-service; distributed training (PyTorch , FSDP), MLflow, and checkpoint management * Design secure, isolated sandbox environments for AI agents and ephemeral execution contexts for untrusted code (gVisor, Kata, Firecracker) * Architect serverless/Lambda-style ephemeral workloads, scale-to-zero, event-driven compute (Knative, KEDA, Firecracker microVMs) * Integrate GPU compute, training, inference, and sandboxes as first-class services in the internal cloud platform * Manage Kubernetes cluster lifecycle at fleet scale using Cluster API (CAPI) provisioning, upgrades, scaling, and decommissioning ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)