> Markdown version of [/jobs/ext/1465937-machine-learning-engineer-ai-foundation](https://www.wearedevelopers.com/jobs/ext/1465937-machine-learning-engineer-ai-foundation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - AI Foundation - **Company:** XPeng Inc. - **Location:** Santa Clara, United States - **Experience:** Expert - **Salary:** $174,720.0 - $295,680.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Nvidia CUDA, Computer Programming, Microprocessors, Python (Programming Language), Graphics Processing Unit (GPU), Pytorch, Large Language Models, TensorRT - **Published:** July 28, 2026 - **Apply:** https://www.dice.com/job-detail/9769261c-4f77-4385-bbe4-90c7e9de2126 ## About the Role * Master in CS/CE/EE, or equivalent, with 3 + years of industry experience. * Good knowledge of PyTorch. * Knowledge of transformer architecture and ways to accelerate the training and inference of transformer models. Preferred Skill Requirements: * Previous experience in the autonomous driving industry. * Knowledge of Torchscript and Nvidia TensorRT. * Strong programming skills in Python and C++ * Familiarity with GPU CPU, NPU, DSP architecture. * Deep understanding of memory bandwidth, compute bottlenecks, and hardware-aware model optimization * Being efficiently in solving complex problems collaboratively on larger teams ## Description * Optimize transformer-based LLMs for low-latency and high-throughput inference. * Optimize kernels and model graphs using tools like CUDA, Triton, and custom fused operators. * Implement and benchmark (Quantization, Knowledge distillation, structured and unstructured pruning, KV-cache optimization, etc.). * Deploy optimized models across GPUs, CPUs, and edge acceleators. * Contribute to internal tooling and documentation for model optimization flows. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts)