Senior Distributed Software Engineer, Golang - DGX Cloud

NVIDIA Corporation
Santa Clara, CA, United States
1 day ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$184,000.0 - $287,500.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Computing Platforms Cloud Computing Software Debugging Distributed Data Store Distributed Systems Open Source Technology Remote Direct Memory Access Cloud Services Software Engineering Enterprise Software Applications Deep Learning
+4 more
Build Management Kubernetes Information Technology Golang

Job description

NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML. In this role you will work closely with the broader NVIDIA team to design and build a reliable, scalable, and efficient storage-as-a-service tailored to AI applications that can be deployed anywhere and scale without limitations. This service supports the whole NVIDIA critical business from graphics drivers to autonomous vehicles to deep learning frameworks. To achieve this goal, we are looking for an engineer with a deep understanding of distributed systems, outstanding design skills, and a track record in building and delivering large-scale distributed services.

What you will be doing:

  • Leading the overall architecture and design of our distributed storage service optimized for AI/ML
  • Develop and maintain distributed, robust and scalable Go programs deployed to state of the art open-source ecosystems, including Kubernetes.
  • Develop and maintain user-space applications, containers, Go-bindings, and CLI tools.
  • Building features for a distributed storage service to enhance availability and reliability for large-scale deployments
  • Engaging and collaborating with NVIDIA Research, Computing, Product teams, cross-functional teams, and external customers to deliver Cloud services.
  • Automating distributed storage service end-to-end, including deployment, management, and monitoring

Requirements

  • Bachelor’s of Science in Computer Science, or related field (or equivalent experience) with 8+ years of industry experience
  • Strong background in developing distributed systems involving Golang, Kubernetes, and Cloud Service Provider integrations
  • Strong track record of delivering distributed services in a variety of distributed computing environments
  • Experience in implementing storage services and interfaces to ensure scalable, high-performance, and reliable solutions
  • History of ownership of product delivery from inception to support
  • Experience developing and maintaining enterprise software. Experience deploying, managing, and debugging applications in a Kubernetes environment
  • Great communication and presentation skills

Ways to stand out from the crowd:

  • You have architected, built, and deployed a distributed service that runs on large-scale clusters, multi-petabyte to exabyte in size, with millions of users
  • You have owned responsibility for all lifecycle stages of software development and delivery.
  • Passionate about innovating and investing in groundbreaking technologies and interested in working with accelerated Computing environments such as GPU Direct Storage, DPU, and RDMA. You are skilled in building and delivering cloud services, with a specific focus on distributed systems

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

About the company

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. Our invention serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is seeking exceptional individuals like you to help us drive the next wave of artificial intelligence.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:25 min

Distinguishing artificial intelligence from deep learning

Sam Witteveen · Coffee With Developers

1:14 min

Contrasting legacy supercomputers with modern AI clustered infrastructure

Thomas Schmidt Thomas Schmidt · World Congress 2024

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

2:08 min

History and scale of NVIDIA GPU computing

Paul Graham Paul Graham · LIVE

Videos

See all

Related articles

See all