AI Training Performance Engineer
Helix
San Jose, United States
1 day ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on startup.jobs
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$200,000.0
Working hours
Regular working hours
Job source
Tech stack
C++ (Programming Language)
Profiling
Nvidia CUDA
Extract Transform Load (ETL)
Software Debugging
Fault Tolerance
InfiniBand
Python (Programming Language)
Regression Analysis
Open Source Technology
Performance Tuning
Remote Direct Memory Access
+5 more
Graphics Processing Unit (GPU)
Pytorch
Information Technology
Performance Monitor
Machine Learning Operations
Job description
- Optimize training performance for a 100B+ parameter models across 100k+ GPUs.
- Collaborate with the broader team on accelerator choice, cluster topology, scheduling, and hardware procurement decisions to inform future scaling.
- Write and optimize custom kernels (Triton/CUDA)
- Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis across training jobs
- Optimize data loading and preprocessing pipelines so I/O never gates the accelerators
- Improve checkpointing, fault tolerance, and elastic restart so large jobs recover quickly from node failures without losing significant wall-clock time
- Partner with researchers to co-design model architectures and training recipes that are performant at scale (e.g., activation checkpointing strategies, mixed precision, sequence packing)
- Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators
- Build and extend agentic systems that automatically generate, benchmark, and iterate on custom kernels
- Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware) for fit with our training workloads, and lead proof-of-concept ports/benchmarks
- Explore different model/data parallelisms (FSDP, context parallel, expert parallel, etc.) to determine optimal configuration per model size.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer/Electrical Engineering, or a related field
- 3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects
- Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy)
- Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ability to translate traces into concrete optimizations
- Solid grasp of collective communication (NCCL) and modern networking concepts (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement).
- Strong Python and CUDA/C++ skills; comfortable reading and modifying framework internals
- Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence)
- Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) and using them to drive optimization priorities, * Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration
- Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.)
- Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management.
Benefits & conditions
The US base salary range for this full-time position is between $200,000 - $400,000 annually., The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on startup.jobs
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
ER
Erin Rifkin
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
about 1 year ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago
CS
Christina Schaireiter
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
3 months ago
DC
Daniel Cranney
Stephan Gillich - Bringing AI Everywhere
almost 2 years ago
BB
Benedikt Bischof
MLOps – What’s the deal behind it?
almost 4 years ago