> Markdown version of [/jobs/ext/2105091-staff-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/2105091-staff-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Machine Learning Engineer - **Company:** XPeng Inc. - **Location:** Santa Clara, United States - **Experience:** Experienced - **Salary:** $215,280.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, Program Optimization, Python (Programming Language), Software Engineering, Pytorch, Delivery Pipeline, Large Language Models, Deep Learning, ONNX (Open Neural Network Exchange) Format, TensorRT - **Published:** August 18, 2026 - **Apply:** https://www.dice.com/job-detail/d0622af4-6a40-423a-9fe3-7b40b1ba156c ## About the Role * Master in CS/CE/EE, or equivalent, with 3-5 years of industry experience. * Strong understanding of Transformer architectures and LLM inference. * Hands-on experience quantizing or deploying deep learning models in production. * Proficiency with PyTorch and at least one inference or compilation stack. * Strong Python programming and software engineering skills. * Ability to work effectively across research, systems, infrastructure, and product teams. * Excellent communication and problem-solving skills, with the ability to thrive in a fast-paced and collaborative environment., * Experience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization. * Experience with AWQ, GPTQ, SmoothQuant, or related methods. * Strong numerical analysis and systems engineering skills. * Experience with one or more LLM runtimes, such as TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes. * Experience deploying LLMs on resource-constrained or heterogeneous hardware. * Contributions to model optimization, inference, compiler, or serving projects. * Publications at NeurIPS, ICML, ICLR, ACL, or related conferences. ## Description * Develop VLA inference models, ensure numerical consistency with training models, and productionize LLM quantization methods, including PTQ, QAT, mixed-precision inference, INT8, FP4, and lower-bit techniques. * Develop production-quality Python code with strong testing, observability, reproducibility, and failure handling. * Build robust model export, calibration, benchmarking, validation, and deployment pipelines. * Engage early with the VLA model research team to establish performance estimates and prove model feasibility. * Curate evaluation datasets and establish a comprehensive metric suite to systematically benchmark VLA performance. * Analyze numerical errors, accuracy regressions, and performance trade-offs. * Develop PTQ and QAT orchestration workflows. * Serve as the primary interface with field-testing and simulation teams for issue triage and autonomous driving performance sign-off. * Collaborate with the in-vehicle software team on latency analysis and issue triage. * Collaborate with the training infrastructure team to develop QAT and model distillation. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)