> Markdown version of [/jobs/ext/2737374-software-engineer-inference-amd-gpu-enablement](https://www.wearedevelopers.com/jobs/ext/2737374-software-engineer-inference-amd-gpu-enablement). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Inference - AMD GPU Enablement - **Company:** OpenAI Inc. - **Location:** San Francisco, CA, United States - **Salary:** $295,000.0 - **Contract:** Permanent contract - **Skills:** Computer Clusters, Nvidia CUDA, Software Debugging, Graphics Processing Unit (GPU), Perf (Linux), Performance Monitor - **Published:** September 5, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/85123067/1 ## About the Role * Have experience writing or porting GPU kernels using HIP, CUDA, or Triton, and care deeply about low-level performance. * Are familiar with communication libraries like NCCL/RCCL and understand their role in high-throughput model serving. * Have worked on distributed inference systems and are comfortable scaling models across fleets of accelerators. * Enjoy solving end-to-end performance challenges across hardware, system libraries, and orchestration layers. * Are excited to be part of a small, fast-moving team building new infrastructure from first principles. Nice to Have: * Contributions to open-source libraries like RCCL, Triton, or vLLM. * Experience with GPU performance tools (Nsight, rocprof, perf) and memory/comms profiling. * Prior experience deploying inference on other non-NVIDIA GPU environments. * Knowledge of model/tensor parallelism, mixed precision, and serving 10B+ parameter models. ## Description We're hiring engineers to scale and optimize OpenAI's inference infrastructure across emerging GPU platforms. You'll work across the stack - from low-level kernel performance to high-level distributed execution - and collaborate closely with research, infra, and performance teams to ensure our largest models run smoothly on new hardware. This is a high-impact opportunity to shape OpenAI's multi-platform inference capabilities from the ground up with a particular focus on advancing inference performance on AMD accelerators. In this role, you will: * Own bring-up, correctness and performance of the OpenAI inference stack on AMD hardware. * Integrate internal model-serving infrastructure (e.g., vLLM, Triton) into a variety of GPU-backed systems. * Debug and optimize distributed inference workloads across memory, network, and compute layers. * Validate correctness, performance, and scalability of model execution on large GPU clusters. * Collaborate with partner teams to design and optimize high-performance GPU kernels for accelerators using HIP, Triton, or other performance-focused frameworks. * Collaborate with partner teams to build, integrate and tune collective communication libraries (e.g., RCCL) used to parallelize model execution across many GPUs. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [How to develop an autonomous car end-to-end: Robotic Drive and the mobility revolution](https://www.wearedevelopers.com/videos/22-how-to-develop-an-autonomous-car-end-to-end-robotic-drive-and-the-mobility-revolution) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs](https://www.wearedevelopers.com/magazine/557-dev-digest-157-cuda-in-python-gemini-code-assist-and-back-dooring-llms) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering)