> Markdown version of [/jobs/ext/1345515-software-engineer-generalist-ai-inference](https://www.wearedevelopers.com/jobs/ext/1345515-software-engineer-generalist-ai-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Generalist, AI Inference - **Company:** Tesla Motors - **Location:** Palo Alto, CA, United States - **Salary:** $118,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Artificial Neural Networks, C++ (Programming Language), Computer Clusters, Nvidia CUDA, Python (Programming Language), Machine Learning, OpenCL, Tensorflow, AI Infrastructure, Application Specific Integrated Circuits, Pytorch, Information Technology, Deployment Automation - **Published:** July 19, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/79997703/1 ## About the Role * Proficiency with Python and C++, including modern C++ (14/17/20) * Proficiency with PyTorch or another machine learning framework * Proficiency with training and deploying neural networks for real-world AI * Proficiency with computer systems and computer architecture * Experience with CUDA ## Description Our team productionalizes ML models - we train and deploy large neural networks for efficient inference on compute-constrained edge devices (CPU / GPU / AI ASIC). The nature of this role is multi-disciplinary - you will work at the intersection of machine learning and systems by building the ML frameworks and infrastructure that enable the seamless training, deployment, and inference of all neural networks that run on Autopilot and Optimus. What You'll Do * Build robust AI frameworks to lower neural networks to edge devices * Build robust AI infrastructure to train and fine-tune networks for Autopilot and Optimus on large GPU clusters * Deploy state-of-the-art neural networks on heterogenous compute, including Tesla's in-house AI ASIC, with an aim to maximize network performance while minimizing latency * Collaborate with AI scientists and compiler engineers to effectively compress large models to run in low precision * Design and implement custom GPU kernels (CUDA / OpenCL) for efficient training and post-processing of network outputs ## Related Videos - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Making neural networks portable with ONNX](https://www.wearedevelopers.com/videos/301-making-neural-networks-portable-with-onnx) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)