> Markdown version of [/jobs/ext/558648-software-engineer-inference-infrastructure](https://www.wearedevelopers.com/jobs/ext/558648-software-engineer-inference-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Inference Infrastructure - **Company:** Tesla Motors - **Location:** Palo Alto, CA, United States - **Salary:** $140,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, C++ (Programming Language), Software Debugging, Programming Tools, File Systems, Distributed Systems, Job Scheduling, Python (Programming Language), Pytorch, Delivery Pipeline, Model Validation, Containerization, Bare Metal, Codebase, Hardware Acceleration, Slurm, Data Pipelines - **Published:** June 15, 2026 - **Apply:** https://diversityjobs.com/career/9340628/Software-Engineer-Foundation-Inference-Infrastructure-California-Palo-Alto ## About the Role * Strong backend engineering fundamentals - distributed systems, job orchestration, reliability, and scale * Experience with hardware accelerator infrastructure - TPUs, custom AI chips, or similar; strong understanding of what it means to manage a large fleet of accelerators and keep them healthy/utilized * Familiarity with cluster orchestration - Kubernetes, SLURM, or similar bare metal & containerized environments * Proficiency in Python; familiarity with PyTorch, Go or C++ is a plus * Experience with low-level systems concepts - networking, file systems, process management * Familiarity with ML inference workloads and what makes them fast or slow at scale * Strong ownership mindset - comfortable navigating ambiguous problems, diving into unfamiliar codebases, and driving things to completion without hand-holding * Experience building CI/CD pipelines for hardware-in-the-loop validation,expertise in fleet management or device provisioning at scale, andfamiliarity with gRPC, distributed task queues, or high-throughput data pipelines are nice-to-haves * Exposure to MLIR or compiler toolchains is a nice-to-have (helpful for working with compiler-produced artifacts and understanding the compilation pipeline) ## Description As a member of the Inference Infrastructure team, you will own the systems that power AI model development, validation, and deployment across custom AI hardware at scale. Your work sits at the intersection of large-scale distributed infrastructure and cutting-edge AI hardware - building the platform that Compiler, AI, and Optimus teams depend on to develop & validate models on next-generation chips.This is not a supporting role; from cluster orchestration & hardware fleet management to inference pipelines & developer tooling,you will design and own foundational systemsthat directly determine how fast the org can move from a trained model to a validated, deployed artifact. What You'll Do * Own & scale the AI inference cluster - the physical and software platform that runs AI workloads on custom AI hardware across thousands of boards * Build & improve job scheduling, hardware onboarding, and cluster self-healing systems that keep the fleet running at 95%+ uptime * Design & implement inference pipelines that unify evals, sims, rollouts, and visualizations across AI and Optimus teams * Build developer tooling that makes compiler-produced artifacts easy to run, validate, and debug on real hardware at scale * Contribute to flashing, inventory management, and fleet management infrastructure for different hardware generations * Work closely with Compiler, AI, and Optimus teams to understand their bottlenecks and build infrastructure that removes them ## Related Videos - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Why your codebase lies to AI?](https://www.wearedevelopers.com/videos/100281-why-your-codebase-lies-to-ai) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)