> Markdown version of [/jobs/ext/3053558-machine-learning-engineer-ai-codesign](https://www.wearedevelopers.com/jobs/ext/3053558-machine-learning-engineer-ai-codesign). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer, AI CoDesign - **Company:** Tesla Motors - **Location:** Palo Alto, CA, United States - **Salary:** $176,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Artificial Neural Networks, C++ (Programming Language), Program Optimization, Nvidia CUDA, Software Debugging, Distributed Computing Environment, Data Flow Control, Python (Programming Language), Linux Kernel, Machine Learning, Toolchain, Application Specific Integrated Circuits, Deep Learning, TensorRT - **Published:** September 24, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/86593450/1 ## About the Role * Strong foundationin deep learning, hands-on experience designing, training, and debugging neural network architectures (transformers, convnets, diffusion models, etc.) * ProficiencywithPyTorch(or equivalent framework), including distributed training, customautogradops, and mixed-precision workflows * Proficiencywith Python and C/C++ (modern C++17/20 preferred) * Solid understanding of computer architecture and systems concepts (memory hierarchy, instruction pipelines, accelerator design) * Experience with model optimization techniques: quantization, pruning, knowledge distillation, or neural architecture search * Experience with CUDA or GPU kernel development * Familiarity with ML compiler stacks or model lowering toolchains (e.g., TVM, XLA, MLIR,TensorRT) is a plus ## Description Our team designs, trains, and deploys large-scale neural networks optimized for inference on compute-constrained edge devices (CPU / GPU / custom AI ASIC). This role sits at the intersection of ML modeling and hardware-aware systems engineering - you will architect and trainstate-of-the-artmodels while co-designing them with the underlying silicon and compiler stack to maximize performance. You will drive the full lifecycle from model research and training at scale to quantized, latency-optimized deployment across Tesla's heterogeneouscomputeplatforms. What You'll Do * Design, train, and iterate on neural network architectures for autonomous driving and robotics, with a focus on efficiency-aware model design (architecture search, distillation, pruning, quantization-aware training) * Co-design model architectures with compiler and ASIC teams to exploit hardware-specific capabilities (custom ops, dataflow patterns, memory hierarchy) * Work cross-functionally with teams on non-standard chip architectures (analog/new material designs) * Profile andoptimizeend-to-end inference latency and throughput across heterogeneous compute targets * Implement custom kernels for training, post-processing, or operations not natively supported by frameworks * Collaborate with AI teams on translating modeling breakthroughs into production-ready, hardware-efficient implementations ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Speeding up Web Apps performance with WebAssembly and Emscripten](https://www.wearedevelopers.com/videos/1985-speeding-up-web-apps-performance-with-webassembly-and-emscripten) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)