> Markdown version of [/jobs/ext/2722587-ai-training-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2722587-ai-training-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Training Performance Engineer - **Company:** Helix - **Location:** San Jose, United States - **Experience:** Experienced - **Salary:** $200,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Profiling, Nvidia CUDA, Extract Transform Load (ETL), Software Debugging, Fault Tolerance, InfiniBand, Python (Programming Language), Regression Analysis, Open Source Technology, Performance Tuning, Remote Direct Memory Access, Graphics Processing Unit (GPU), Pytorch, Information Technology, Performance Monitor, Machine Learning Operations - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/helix-ai-engineer-training-performance-figure-2-9004636 ## About the Role * Bachelor's or Master's degree in Computer Science, Computer/Electrical Engineering, or a related field * 3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects * Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy) * Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ability to translate traces into concrete optimizations * Solid grasp of collective communication (NCCL) and modern networking concepts (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement). * Strong Python and CUDA/C++ skills; comfortable reading and modifying framework internals * Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence) * Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) and using them to drive optimization priorities, * Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration * Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.) * Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management. ## Description * Optimize training performance for a 100B+ parameter models across 100k+ GPUs. * Collaborate with the broader team on accelerator choice, cluster topology, scheduling, and hardware procurement decisions to inform future scaling. * Write and optimize custom kernels (Triton/CUDA) * Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis across training jobs * Optimize data loading and preprocessing pipelines so I/O never gates the accelerators * Improve checkpointing, fault tolerance, and elastic restart so large jobs recover quickly from node failures without losing significant wall-clock time * Partner with researchers to co-design model architectures and training recipes that are performant at scale (e.g., activation checkpointing strategies, mixed precision, sequence packing) * Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators * Build and extend agentic systems that automatically generate, benchmark, and iterate on custom kernels * Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware) for fit with our training workloads, and lead proof-of-concept ports/benchmarks * Explore different model/data parallelisms (FSDP, context parallel, expert parallel, etc.) to determine optimal configuration per model size. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [30 Golden Rules of Deep Learning Performance](https://www.wearedevelopers.com/videos/11-30-golden-rules-of-deep-learning-performance) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Enhancing Workload Security in Kubernetes](https://www.wearedevelopers.com/videos/356-enhancing-workload-security-in-kubernetes) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)