Senior Systems Software Engineer - NV Cloud Functions

NVIDIA Corporation
Santa Clara, CA, United States
4 days ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$152,000.0 - $241,500.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Bash Shell Cloud Engineering Encodings Linux Distributed Systems Python (Programming Language) Load Testing Machine Learning Open Source Technology Performance Tuning Reliability Engineering
+15 more
Software Engineering Rust (Programming Language) Graphics Processing Unit (GPU) High Performance Computing System Availability Large Language Models Deep Learning AI Platforms Kubernetes Information Technology Data Analytics Slurm TensorRT Serverless Computing Golang

Job description

NVIDIA Cloud Functions team is looking for a motivated, product-minded AI/ML Engineer with domain expertise in AI platform engineering at scale. Our team builds and operates a serverless deployment platform for enabling AI applications. Our product enables and scales AI inferencing workloads using globally distributed orchestration of workloads on GPU-backed cloud-agnostic Kubernetes clusters. You will be working with a team of passionate and skilled engineers that are continuously innovating at the speed of light to provide the best product possible, for both external customers and internal NVIDIA teams. We are looking for someone to join us at the forefront of defining cloud engineering paradigms for AI at scale.

What You’ll be Doing:

  • Becoming a trusted subject matter expert by understanding user challenges and constraints. Translate this into product requirements and solutions, accelerating delivery of AI models and inference hosted on the NVCF platform.
  • Leading implementation of key features. Conducting user-acceptance testing, load testing and performance evaluations. Emphasis on customer experience, performance optimization and platform reliability.
  • Mentoring and embedding with other engineering teams building products on top of our platform on best practices for AI/ML workloads at scale with excellent performance, including ML reliability engineering at scale.
  • Producing reference architectures to guide customer use cases, applying the newest NVCF product features with the latest AI technologies.
  • Shepherding customer issues to resolution and providing timely warning of issues and risks.
  • Evaluating new and innovative technologies and tooling as the AI-at-scale landscape evolves to ensure we have a competitive product and forward-looking roadmap.

Requirements

  • Masters, PhD, or equivalent experience in Computer Science, Artificial Intelligence, Applied Math, or related field
  • At least 2 years work experience with Python, Rust, Golang, Linux or Bash.
  • Experience in Deep Learning and Machine Learning; expertise in using AI/DL frameworks and inferencing software such as SGLang, vLLM, TensorRT-LLM, or Dynamo.
  • Knowledge of CPU and GPU architecture.
  • Excellent interpersonal skills including ability to explain sophisticated technical topics to non-experts.
  • Experience in the design, implementation, and release of AI/ML products to market. A flexible technologist familiar with all aspects of the software development lifecycle.

Ways to stand out from the crowd:

  • Demonstrate a strong desire to share knowledge with clients, partners and co-workers, able to show this through previous work.
  • Demonstrate expertise through projects or Open Source contributions in HPC, Data Analytics, Machine Learning, Deep Learning, Cloud Native Projects, Kubernetes, Slurm, or enabling GPU workloads.
  • Show a willingness and ability to dig into unfamiliar territories to tackle complex problems through examples in previous work.
  • Prior experience in building distributed systems.

Benefits & conditions

NVIDIA offers highly competitive salaries and a comprehensive benefits package. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most brilliant and talented people in the world working for us. Are you a creative engineer with a drive for advancing the state of AI and bringing it to the cloud? If you love to tackle problems and advocate for continuous, innovative improvement, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all