> Markdown version of [/jobs/ext/2840281-staff-ml-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2840281-staff-ml-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff ML Performance Engineer - **Company:** WAYVE LLC - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Salary:** $336,400.0 - **Contract:** Permanent contract - **Skills:** Profiling, Nvidia CUDA, Distributed Systems, Python (Programming Language), Machine Learning, Performance Tuning, Information Technology, Machine Learning Operations - **Published:** September 11, 2026 - **Apply:** https://www.thejobnetwork.com/job/99ad485c-aafc-4d37-9dce-565863c1c109/staff-ml-performance-engineer-training-efficiency ## About the Role - 10+ years of industry experience driving performance engineering across ML systems, GPU compute infrastructure, distributed platforms or similar field. - Experience optimizing large scale jobs on GPU compute clusters. - Experience in working in platform teams and working with research teams. - Experience in writing, reporting, and tracking performance benchmarks in an open and accessible way. - Ability to write high quality, well-structured and tested Python code - BS or MS in Machine Learning, Computer Science, Engineering, or a related technical discipline or equivalent experience **Desirable** - Experience working with concurrent, parallel and distributed computing. - Experience using NVIDIA NSight Systems or other system profilers. - Experience implementing GPU kernels (CUDA, Triton, etc). - Knowledge of computing fundamentals - what makes code fast, secure and reliable. ## Description We are looking for a Staff ML Performance Engineer to join our Training Tech team working on optimizing large scale ML jobs to enable scaling our models to the next order of magnitude. A successful candidate will increase efficiency of training and inference workloads in order to allow Wayve to train larger models faster. Key responsibilities: - Profile ML workloads to identify their bottlenecks, e.g. using NVIDIA Nsight Systems - Design and implement efficiency improvements to maximize MFU and throughput, e.g. parallelism, model compilation, mixed precision - Design and implement observability tools to identify bottlenecks and drive performance improvements, e.g. to track MFU, throughput, latency, etc - Design and implement benchmarking tools, e.g. to track efficiency gains or regressions - Collaborate closely with Research teams to integrate training efficiency improvements and create a culture of performance optimization ## **About you** In order to set you up for success in this role, we're looking for the following skills and experience. ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/1521-accelerating-python-on-gpus) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)