Senior System Software Engineer - AI Data Platform - Inference Factory Optimization

NVIDIA Ltd.
Santa Clara, CA, United States
1 day ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Business Analytics Applications Data Analysis Automation of Tests C++ (Programming Language) Profiling Computer Programming Computer Engineering Continuous Delivery Continuous Integration Data Centers Data Infrastructure
+14 more
Software Debugging Device Drivers Distributed Systems Memory Management Python (Programming Language) Machine Learning Network Protocols Performance Tuning Software Engineering Software Systems System Software Reliability of Systems Information Technology Docker

Job description

  • Develop efficient infrastructure and tools for automating complex software processes.
  • Drive Performance Optimization: Implement advanced test harnesses, benchmarking frameworks, and analytical tools to rigorously characterize and optimize the performance and efficiency of our software and hardware platforms.
  • Apply deep knowledge of operating systems, kernel internals, device drivers, memory management, storage, networking, and high-speed interconnects to build and troubleshoot highly performant systems.
  • Work with engineering teams to understand needs, define requirements, and deliver efficient solutions.
  • Set performance goals, monitor feedback, analyze data, and make continuous improvements for system reliability.
  • Influence Technical Strategy: Contribute to defining technical strategies and roadmaps for our platform automation initiatives, ensuring alignment with company-wide goals and standard methodologies.

Requirements

  • Bachelor’s or equivalent experience in Computer Science, Computer Engineering, or a related technical field, or Master’s degree or equivalent experience in a similar field.
  • 5+ years of industry experience in software development, focusing on infrastructure, distributed systems, automation, and/or performance engineering.
  • Expertise in System-Level Programming: Proven ability to develop robust tools and automation using programming languages such as C++, Python, or Go.
  • Deep Understanding of System Software: Experience with operating system internals, device drivers, memory management, and debugging performance issues in complex compute applications.
  • Distributed Systems: Experience in designing, building, and operating large-scale distributed systems, with knowledge of networking protocols, cluster management, and high-performance interconnects.
  • Automation and CI/CD Proficiency: Experience building and maintaining automated testing, benchmarking, and continuous integration/continuous deployment pipelines.
  • Problem-Solving and Analytical Skills: Outstanding analytical, problem-solving, and debugging skills, with a track record of resolving complex technical challenges.
  • Collaboration and Communication: Excellent interpersonal and communication skills, with the ability to articulate complex technical concepts to diverse audiences and collaborate effectively across teams.

Ways to stand out from the crowd

  • Experience optimizing performance for AI/Machine Learning workloads, especially inference applications, on diverse hardware platforms.
  • Prior experience building or contributing to large-scale compute infrastructure solutions in cloud environments or on-premises data centers.
  • Experience with containerization and orchestration technologies, such as Docker and Kubernetes.
  • Familiarity with performance profiling tools and methodologies for hardware and software systems.
  • Track record of driving significant efficiency gains or architectural improvements in large-scale systems.

Benefits & conditions

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

1:58 min

Verifying hardware access and exploring AI inference scaling

Piotr Zaniewski Piotr Zaniewski · World Congress 2026 Europe

Videos

See all

Related articles

See all