Senior Manager, Performance Engineering - Kernel and Software Platforms
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview In this role, you guide an engineering team to maintain peak product performance across the full hardware lifecycle. You bridge design-time predictions with real-world execution, spanning pre-silicon to production hardware and post-release support. You shape performance expectations, curate workloads, and align with CUDA release timelines while advancing automation with AI tooling. This is a hands-on leadership position shaping how NVIDIA delivers high-performance DL architectures and AI computing platforms. Compensation / Benefits * Oversee end-to-end performance tracking from pre-silicon through production hardware and maintenance * Build and refine theoretical and empirical performance models and correlate predictions with real hardware telemetry * Evaluate DSL-to-IR performance flows from high-level DSLs (e.g., Triton, PyTorch) to target code across platform maturity * Create and maintain stress-test workloads and testlists to detect regressions and validate releases * Coordinate CUDA release cadence with compiler, architecture, and platform software teams * Drive workflow automation with AI tooling (Claude, Codex, agentic LLM workflows) to streamline telemetry analysis and reporting pipelines Tasks * 10+ years in systems/software performance engineering, benchmarking, or related fields * 5+ years leading or managing technical engineering teams * Experience tracking performance through full hardware lifecycle from design-time to production hardware * Proven ability to build and correlate performance models with hardware telemetry without daily low-level kernel development * Understanding of kernel compilation pipelines and compiler flows (DSL -> IR -> target code) and their impact on execution efficiency * Experience developing workload testlists and aligning performance delivery with major software release cycles (CUDA cadence) * Proficiency in Python and practical use of generative AI APIs/models (Codex, Claude) for automation and analytics Key requirements * equity * benefits * hybrid work arrangement * career growth * competitive salary * relocation assistance
Requirements
aay_ * Coordinate CUDA release cadence with compiler, architecture, and platform software teams * Drive workflow automation with AI tooling (Claude, Codex, agentic LLM workflows) to streamline telemetry analysis and reporting pipelines Tasks * 10+ years in systems/software performance engineering, benchmarking, or related fields * 5+ years leading or managing technical engineering teams * Experience tracking performance through full hardware lifecycle from design-time to production hardware * Proven ability to build and correlate performance models with hardware telemetry without daily low-level kernel development * Understanding of kernel compilation pipelines and compiler flows (DSL -> IR -> target code) and their impact on execution efficiency * Experience developing workload testlists and aligning performance delivery with major software release cycles (CUDA cadence) * Proficiency in Python and practical use of generative AI APIs/models (Codex, Claude) for automation and analytics Key a
Benefits & conditions
Engineering * equity * benefits * hybrid work arrangement * career growth * competitive salary * relocation assistance
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Highest Paying Tech Companies for Developers
MLOps – What’s the deal behind it?