Sr. Software Development Engineer - SGLang and Inference Stack
Role details
Job location
Tech stack
Job description
As a core member of the team, you will play a pivotal role inoptimizingand developing deep learning frameworks for AMD GPUs. Your work will be instrumental in enhancing GPU kernel performance, accelerating deep learning models, and enabling RL training and SOTA LLM and Multimodal inference at scale across multi-GPU and multi-node systems. You will collaborate across internal GPU software teams and engage with open-source communities to integrateand optimize cutting-edgecompiler technologies and drive upstream contributions thatbenefitAMD's AI software ecosystem., * OptimizeDeep Learning Frameworks: Enhance performance of frameworks like TensorFlow,PyTorch, andSGLangon AMD GPUs via upstream contributions in open-source repositories.
-
Develop and Optimize Deep Learning Models: Profile, analyze, code change and tune large-scale training and inference models foroptimalperformance on AMD hardware.Day-0 supports to many SOTA models, DeepSeek 3.2, Kimi K2.5, etc.
-
GPU Kernel Development: Design, implement, andoptimizehigh-performance GPU kernels using HIP, Triton, TileLang or other DSLs for AI operator efficiency.
-
Collaborate with GPU Library and Compiler Teams: Work closely with internal compiler and GPU math library teams to integrate, optimize and align kernel-level optimizations with full-stack performance goals.Initiate and help with different level codegen optimizations.
-
Contribute toSGLangDevelopment: Support optimization, feature development, and scaling of theSGLangframework across AMD GPU platforms for LLM, multimodal serving and RL-training.
-
Distributed System Optimization: Tune and scale performance across both multi-GPU (scale-up) and multi-node (scale-out) environments, including inference parallelism, prefill-decode disaggregation, Wide-EP and collective communication strategies.
-
Graph Compiler Integration: Integrate andoptimizeruntime execution through graph compilers such as XLA,TorchDynamo, or custom pipelines.
-
Open-Source Collaboration: Partner with external maintainers to understand framework needs, propose optimizations, and upstream contributions effectively.
-
Apply Engineering Best Practices: Leverage modern software engineering practices in debugging, profiling, test-driven development, and CI/CD integration., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here.
Requirements
Skilled engineer with strong technical and analyticalexpertisein GPGPU C++, Triton, TileLang or DSL development within Linux environments. The ideal candidate will thrive in both collaborative team settings and independent work, with the ability to define goals, manage development efforts, and deliver high-quality solutions. Strong problem-solving skills, a proactive approach, and a keen understanding of software engineering best practices are essential., * Strong Programming Skills:Proficient in C++ and/or Python (PyTorch, Triton, TileLang), withdemonstratedability to code, debug, profile, and optimize performance-critical code.
-
SGLangand LLM Optimization:Hands-on experience withSGLangor similar LLM inference frameworks is highly preferred.
-
Compiler and GPU Architecture Knowledge:Background in compiler design or familiarity with technologies like LLVM, MLIR, orROCmis a plus.
-
Heterogeneous System Workloads:Experience running and scaling workloads on large-scale, heterogeneous clusters (CPU + GPU) using distributed training or inference strategies.
-
AI Framework Integration:Experience contributing to or integrating optimizations into deep learning frameworks such asPyTorch, SGLang, vLLM, Slime, VeRL
-
GPGPU Computing:Working knowledge of HIP, CUDA, Triton, TileLang or other GPU programming models; experience with GCN/CDNA architecture preferred., * Bachelor's and/orMaster'sDegreein Computer Science, Computer Engineering, Electrical Engineering, Physics or a related field.