> Markdown version of [/jobs/ext/2281151-principal-software-engineer-e2e-performance-and-goodput-csp-engagements](https://www.wearedevelopers.com/jobs/ext/2281151-principal-software-engineer-e2e-performance-and-goodput-csp-engagements). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Engineer, E2E Performance and Goodput - CSP Engagements - **Company:** NVIDIA Corporation - **Location:** Santa Clara, CA, United States - **Salary:** $272,000.0 - **Contract:** Permanent contract - **Skills:** Profiling, Nvidia CUDA, Computer Engineering, Memory Management, Firmware, Python (Programming Language), Open Source Technology, Pattern Recognition, Strategies of Testing, Large Language Models, Pandas, Information Technology, Machine Learning Operations, TensorRT - **Published:** August 28, 2026 - **Apply:** https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Software-Engineer--E2E-Performance-and-Goodput---CSP-Engagements_JR2020321 ## About the Role * 15+ years of experience in systems performance engineering, ideally in GPU/HPC/ML infrastructure. BS or MS in Computer Science, Computer Engineering, or related field (or equivalent experience) * Proficiency in GPU workload profiling: nsight systems, nsight compute, DCGM metrics, or equivalent instrumentation * Understanding of distributed training performance dynamics: computation/communication overlap, pipeline bubbles, memory bandwidth utilization, collective efficiency * Statistical methods for performance analysis: regression detection, confidence intervals, A/B comparison at scale * Understanding of how the full software stack impacts performance: driver overhead, collective algorithm selection, memory allocation, scheduling, firmware power management * Strong data analysis and visualization skills (Python, pandas, dashboards). Customer obsession - genuine passion for understanding why customers aren't achieving expected performance and driving solutions * Ability to communicate performance findings to both deep technical audiences and executive leadership * Demonstrated success influencing multiple engineering teams to prioritize performance improvements Ways to stand out from the crowd: * Experience profiling and optimizing distributed training at 1000+ GPU scale (Megatron-LM, DeepSpeed, FSDP) * Background in ML infrastructure performance at a CSP/hyperscaler * Familiarity with NVIDIA platforms (DGX, HGX, NVLink topology) and profiling tools * Experience building automated performance regression detection systems for production environments * Understanding of inference workload performance dynamics (vLLM, TensorRT-LLM, SGLang, continuous batching) ## Description We're looking for a Principal Engineer to join our CSP Engagements team as the technical focal point for end-to-end performance, working directly with engineering teams of key CSP/hyperscale customers to ensure they achieve various performance targets on NVIDIA platforms. In this role, you will augment NVIDIA's performance and benchmark teams with a dedicated CSP-facing focus. You will drive work streams with CSP engineering teams to build shared understanding of platform performance characteristics, gather and incorporate their workload-specific feedback into NVIDIA's optimization priorities, and validate that performance targets are met in customer-representative configurations. Your cross-CSP visibility enables you to identify patterns and drive systemic improvements in documentation, configuration guidance, and tooling. What you'll be doing: * Drive performance characterization work streams with engineering teams of key CSP/hyperscale customers - ensuring they understand platform performance expectations, profiling methodology, and tuning options for their specific workloads * Gather and synthesize CSP performance feedback - identify gaps between expected and actual throughput, and champion optimization priorities back into NVIDIA's CUDA, NCCL, driver, and firmware teams * Ensure key open-source performance and stress tools (e.g., STREAM, GPU Burn, GPU BLAST) are updated and validated for the latest NVIDIA rack-scale systems, GPU architectures, and CPU platforms - so customers and internal teams have reliable baseline measurements from day one * Work closely with CSPs to ensure their own performance and validation tooling reflects the latest GPU capabilities, memory hierarchy changes, and platform-specific tuning parameters * Conduct cross-CSP performance comparison and pattern analysis - identify configuration, software, or workload differences that explain performance gaps between deployments * Collaborate with CSPs to ensure performance-related integration work (profiling infrastructure, benchmark harnesses, config validation) is ready ahead of deployment milestones * Define test strategies and tooling requirements for performance validation - both for NVIDIA internal certification and customer acceptance ## Related Videos - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers)