Performance Engineer (GPU)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
- As a GPU Performance Engineer, you’ll architect and implement the foundational systems that power Claude and push the frontiers of what’s possible with large language models
- You’ll be responsible for maximizing GPU utilization and performance at unprecedented scale, developing cutting-edge optimizations that directly enable new model capabilities and dramatically improve inference efficiency
- Working at the intersection of hardware and software, you’ll implement state-of-the-art techniques from custom kernel development to distributed system architectures
- Your work will span the entire stack-from low-level tensor core optimizations to orchestrating thousands of GPUs in perfect synchronization
- Co-design attention mechanisms and algorithms for next-generation hardware architectures
- Develop custom kernels for emerging quantization formats and mixed-precision techniques
- Design distributed communication strategies for multi-node GPU clusters
- Optimize end-to-end training and inference pipelines for frontier language models
- Build performance modeling frameworks to predict and optimize GPU utilization
- Implement kernel fusion strategies to minimize memory bandwidth bottlenecks
- Create resilient systems for planet-scale distributed training infrastructure
- Profile and eliminate performance bottlenecks in production serving infrastructure
- Partner with hardware vendors to influence future accelerator capabilities and software stacks
Requirements
Strong candidates will have a track record of delivering transformative GPU performance improvements in production ML systems and will be excited to shape the future of AI infrastructure alongside world-class researchers and engineersHave deep experience with GPU programming and optimization at scaleCare about the societal impacts of your workCan navigate complex systems from hardware interfaces to high-level ML frameworksAre impact-driven, passionate about delivering measurable performance breakthroughsEnjoy collaborative problem-solving and pair programmingThrive in ambiguous environments where you define the path forwardWant to work on state-of-the-art language models with real-world impactEducation requirements: We require at least a Bachelor’s degree in a related field or equivalent experienceGPU Kernel Development: CUDA, Triton, CUTLASS, Flash Attention, tensor core optimizationML Compilers & Frameworks: PyTorch/JAX internals, pile, XLA, custom operatorsPerformance Engineering: Kernel fusion, memory bandwidth optimization, profiling with NsightDistributed Systems: NCCL, NVLink, collective communication, model parallelismLow-Precision: INT8/FP8 quantization, mixed-precision techniquesProduction Systems: Large-scale training infrastructure, fault tolerance, cluster orchestration
Benefits & conditions
- Comprehensive health, dental, and vision insurance for you and your dependents
- Inclusive fertility benefits via Carrot Fertility
- 22 weeks of paid parental leave
- Flexible paid time off and absence policies
- Mental health support for you and your dependents
- Competitive salary and equity packages
- Optional equity donation matching at a 1:1 ratio, up to 25% of your equity grant
- Retirement plans with competitive matching
- Life and income protection plans
- $500/month flexible wellness and time saver stipend
- Commuter benefits
- Annual education stipend
- Home office stipends
- Relocation support for those moving for Anthropic
- Daily meals and snacks in the office
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…