AI Training Performance Engineer

Helix
San Jose, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$200,000.0
Working hours
Regular working hours
Job source

Tech stack

C++ (Programming Language) Profiling Nvidia CUDA Extract Transform Load (ETL) Software Debugging Fault Tolerance InfiniBand Python (Programming Language) Regression Analysis Open Source Technology Performance Tuning Remote Direct Memory Access
+5 more
Graphics Processing Unit (GPU) Pytorch Information Technology Performance Monitor Machine Learning Operations

Job description

  • Optimize training performance for a 100B+ parameter models across 100k+ GPUs.
  • Collaborate with the broader team on accelerator choice, cluster topology, scheduling, and hardware procurement decisions to inform future scaling.
  • Write and optimize custom kernels (Triton/CUDA)
  • Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis across training jobs
  • Optimize data loading and preprocessing pipelines so I/O never gates the accelerators
  • Improve checkpointing, fault tolerance, and elastic restart so large jobs recover quickly from node failures without losing significant wall-clock time
  • Partner with researchers to co-design model architectures and training recipes that are performant at scale (e.g., activation checkpointing strategies, mixed precision, sequence packing)
  • Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators
  • Build and extend agentic systems that automatically generate, benchmark, and iterate on custom kernels
  • Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware) for fit with our training workloads, and lead proof-of-concept ports/benchmarks
  • Explore different model/data parallelisms (FSDP, context parallel, expert parallel, etc.) to determine optimal configuration per model size.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Computer/Electrical Engineering, or a related field
  • 3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects
  • Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy)
  • Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ability to translate traces into concrete optimizations
  • Solid grasp of collective communication (NCCL) and modern networking concepts (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement).
  • Strong Python and CUDA/C++ skills; comfortable reading and modifying framework internals
  • Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence)
  • Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) and using them to drive optimization priorities, * Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration
  • Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.)
  • Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management.

Benefits & conditions

The US base salary range for this full-time position is between $200,000 - $400,000 annually., The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

4:15 min

Diagnosing training performance bottlenecks with visual profiling tools

Anirudh Koul · LIVE

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

Videos

See all

Related articles

See all