> Markdown version of [/jobs/ext/2283986-senior-gpu-performance-software-engineer](https://www.wearedevelopers.com/jobs/ext/2283986-senior-gpu-performance-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior GPU Performance Software Engineer - **Company:** Intel Corporation - **Location:** Hillsboro, OR, United States (Remote available) - **Experience:** Expert - **Salary:** $195,200.0 - $275,580.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Profiling, Nvidia CUDA, Computer Programming, Computer Engineering, Microprocessors, Linux Kernel, Machine Learning, OpenMP, Open Source Technology, OpenCL, Performance Tuning, Tensorflow, Software Engineering, Multithreading, Graphics Processing Unit (GPU), Pytorch, Deep Learning, Parallel Computation, Information Technology, ONNX (Open Neural Network Exchange) Format - **Published:** August 28, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3367766773&tx=KL707PFK&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role To be successful in this role, you should demonstrate the following professional traits: * A strong ownership mindset - you take initiative on complex, ambiguous technical problems and drive them to resolution * A collaborative approach - you work effectively across hardware, compiler, and framework teams to align on shared technical goals * A performance-driven curiosity - you are motivated by squeezing every cycle out of hardware and continuously seek deeper understanding of low-level systems, * Education: BSc, MSc, or PhD in Computer Science, Computer Engineering, Mathematics, Physics, or a highly technical related field * Core Language: 5+ years of professional software development experience with expert-level modern C++ Performance Optimizations: 2+ years of hands-on experience in programming and kernel optimization on GPUs (via SYCL/DPC++, OpenCL, CUDA, or HIP), or at least 5+ years of similar low-level performance optimization experience on CPUs * Hardware Architecture: Strong foundations in computer architecture, cache hierarchies, memory subsystems, and parallel programming paradigms (e.g., multi-threading, SIMD/vectorization), * Math Libraries:Experience developing high-performance math libraries (e.g., GEMM, convolution, reduction, or FFT kernels) * Low-Level Tuning:Hands-on experience with GPU assembly-level tuning or compiler optimization * Parallel APIs:Familiarity with parallel programming APIs such as OpenMP or oneTBB * AI Workload Context:Basic understanding of deep learning primitives (e.g., forward/backward passes) to understand how library code is utilized by upstream frameworks ## Description The Software and AI (SAI) organization is seeking a highly skilled software engineer to contribute to the development and low-level optimization of oneDNN , a complex, cross-platform, open-source performance library that serves as the foundation for deep learning applications ( github.com/uxlfoundation/oneDNN ). Please Note: This is a low-level software engineering and hardware-acceleration role. It does not involve building, training, or tuning machine learning models. Instead, you will focus on developing highly optimized math primitives, parallel algorithms, and GPU kernels that power industry-leading AI frameworks (such as OpenVINO, TensorFlow, PyTorch, and ONNX Runtime) on Intel hardware. Key Responsibilities Kernel Development and Architecture * Develop high-performance GEMM, convolution, and attention kernels for AI workloads * Design scalable JIT and codegen infrastructure for GPU kernel generation Low-Level Optimization * Implement fusion and memory-traffic optimizations to maximize hardware utilization * Optimize mixed-precision and quantized execution paths (e.g., BF16, FP16, INT8, FP8, FP4, etc.) Performance Modeling and Profiling * Build analytical and empirical performance models for kernel dispatch and tuning * Profile and eliminate performance bottlenecks across oneDNN GPU primitives and runtime paths Hardware and Software Co-Design * Co-design GPU primitives and kernel architectures for next-generation Intel GPUs * Partner with hardware and compiler teams to shape future accelerator capabilities and software stacks Infrastructure and Validation * Improve validation, benchmarking, and CI infrastructure for performance-critical GPU workloads, This role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. * Job posting details (such as work model, location or time type) are subject to change. ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/1521-accelerating-python-on-gpus) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/1112-accelerating-python-on-gpus) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023)