Senior Manager, Performance Engineering - Kernel and Software Platforms

NVIDIA Ltd.
Santa Clara, CA, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Computing Platforms Nvidia CUDA Python (Programming Language) Linux Kernel Software Engineering Pytorch Large Language Models Software Performance

Job description

Experteer Overview In this role, you guide an engineering team to maintain peak product performance across the full hardware lifecycle. You bridge design-time predictions with real-world execution, spanning pre-silicon to production hardware and post-release support. You shape performance expectations, curate workloads, and align with CUDA release timelines while advancing automation with AI tooling. This is a hands-on leadership position shaping how NVIDIA delivers high-performance DL architectures and AI computing platforms. Compensation / Benefits * Oversee end-to-end performance tracking from pre-silicon through production hardware and maintenance * Build and refine theoretical and empirical performance models and correlate predictions with real hardware telemetry * Evaluate DSL-to-IR performance flows from high-level DSLs (e.g., Triton, PyTorch) to target code across platform maturity * Create and maintain stress-test workloads and testlists to detect regressions and validate releases * Coordinate CUDA release cadence with compiler, architecture, and platform software teams * Drive workflow automation with AI tooling (Claude, Codex, agentic LLM workflows) to streamline telemetry analysis and reporting pipelines Tasks * 10+ years in systems/software performance engineering, benchmarking, or related fields * 5+ years leading or managing technical engineering teams * Experience tracking performance through full hardware lifecycle from design-time to production hardware * Proven ability to build and correlate performance models with hardware telemetry without daily low-level kernel development * Understanding of kernel compilation pipelines and compiler flows (DSL -> IR -> target code) and their impact on execution efficiency * Experience developing workload testlists and aligning performance delivery with major software release cycles (CUDA cadence) * Proficiency in Python and practical use of generative AI APIs/models (Codex, Claude) for automation and analytics Key requirements * equity * benefits * hybrid work arrangement * career growth * competitive salary * relocation assistance

Requirements

aay_ * Coordinate CUDA release cadence with compiler, architecture, and platform software teams * Drive workflow automation with AI tooling (Claude, Codex, agentic LLM workflows) to streamline telemetry analysis and reporting pipelines Tasks * 10+ years in systems/software performance engineering, benchmarking, or related fields * 5+ years leading or managing technical engineering teams * Experience tracking performance through full hardware lifecycle from design-time to production hardware * Proven ability to build and correlate performance models with hardware telemetry without daily low-level kernel development * Understanding of kernel compilation pipelines and compiler flows (DSL -> IR -> target code) and their impact on execution efficiency * Experience developing workload testlists and aligning performance delivery with major software release cycles (CUDA cadence) * Proficiency in Python and practical use of generative AI APIs/models (Codex, Claude) for automation and analytics Key a

Benefits & conditions

Engineering * equity * benefits * hybrid work arrangement * career growth * competitive salary * relocation assistance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

1:15 min

Overcoming the challenges of modifying Linux kernel code

Ayesha Kaleem · WWC 2023

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · WWC Europe 2026

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

Videos

See all

Related articles

See all