> Markdown version of [/jobs/ext/3564597-performance-engineer-gpu](https://www.wearedevelopers.com/jobs/ext/3564597-performance-engineer-gpu). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Performance Engineer (GPU) - **Company:** Anthropic - **Location:** North Yorkshire, UK (Remote available) - **Contract:** Permanent contract - **Skills:** Adobe Flash, Computer Clusters, Compilers, Profiling, Nvidia CUDA, Distributed Systems, Fault Tolerance, Hardware Interface Design, Linux Kernel, Language Modeling, AI Infrastructure, Graphics Processing Unit (GPU), Pytorch, Delivery Pipeline, Large Language Models, Gpu Programming, Machine Learning Operations, Claude, CUTLASS - **Published:** October 3, 2026 - **Apply:** https://find.jobs/jobs-near-me/apply/ats-redirect/?id=2999451072-2 ## About the Role Strong candidates will have a track record of delivering transformative GPU performance improvements in production ML systems and will be excited to shape the future of AI infrastructure alongside world-class researchers and engineersHave deep experience with GPU programming and optimization at scaleCare about the societal impacts of your workCan navigate complex systems from hardware interfaces to high-level ML frameworksAre impact-driven, passionate about delivering measurable performance breakthroughsEnjoy collaborative problem-solving and pair programmingThrive in ambiguous environments where you define the path forwardWant to work on state-of-the-art language models with real-world impactEducation requirements: We require at least a Bachelor's degree in a related field or equivalent experienceGPU Kernel Development: CUDA, Triton, CUTLASS, Flash Attention, tensor core optimizationML Compilers & Frameworks: PyTorch/JAX internals, pile, XLA, custom operatorsPerformance Engineering: Kernel fusion, memory bandwidth optimization, profiling with NsightDistributed Systems: NCCL, NVLink, collective communication, model parallelismLow-Precision: INT8/FP8 quantization, mixed-precision techniquesProduction Systems: Large-scale training infrastructure, fault tolerance, cluster orchestration ## Description * As a GPU Performance Engineer, you'll architect and implement the foundational systems that power Claude and push the frontiers of what's possible with large language models * You'll be responsible for maximizing GPU utilization and performance at unprecedented scale, developing cutting-edge optimizations that directly enable new model capabilities and dramatically improve inference efficiency * Working at the intersection of hardware and software, you'll implement state-of-the-art techniques from custom kernel development to distributed system architectures * Your work will span the entire stack-from low-level tensor core optimizations to orchestrating thousands of GPUs in perfect synchronization * Co-design attention mechanisms and algorithms for next-generation hardware architectures * Develop custom kernels for emerging quantization formats and mixed-precision techniques * Design distributed communication strategies for multi-node GPU clusters * Optimize end-to-end training and inference pipelines for frontier language models * Build performance modeling frameworks to predict and optimize GPU utilization * Implement kernel fusion strategies to minimize memory bandwidth bottlenecks * Create resilient systems for planet-scale distributed training infrastructure * Profile and eliminate performance bottlenecks in production serving infrastructure * Partner with hardware vendors to influence future accelerator capabilities and software stacks