Senior Software Engineer, Distributed Systems...

NVIDIA Ltd.
Santa Clara, CA, United States
24 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$168,000.0 - $270,250.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Automation of Tests Cloud Computing Computer Engineering Software Debugging Distributed Systems Redis Prometheus Graphics Processing Unit (GPU) Cloud Platform System Data Ingestion Deep Learning
+8 more
Backend Build Management Kubernetes Information Technology Apache Kafka Nim (Programming Language) Docker Microservices

Job description

NVIDIA is building a new category of products, by intersecting our prowess in deep learning and computing, with industry-leading technologies. You will harness groundbreaking technologies, and build a highly efficient factory to power how NVIDIA builds and validates NIMs for inferencing all the way through deployment in heterogeneous hardware and software environments. You will influence and drive technical advances in NVIDIAs workflows and build the infrastructure that strives to accelerate the delivery of every AI model on NVIDIA’s GPUs anywhere. We are looking for technical talent to design and build our factory capabilities, including the underlying infrastructure, pipelines, backends, Docker build, test harness, metrics, performance engineering, log ingestion, and more.

What you’ll be doing:

  • Develop a factory pipeline that will take an AI model in and produce a deployable service that is validated across Cloud, On-prem and Kubernetes environments. With the team, define and deliver rapid iterations on the group’s technical strategies and roadmaps to deliver and improve the NIM factory. You will be designing interfaces, data modeling and schema design, and expanding observability over the factory pipeline and its compute infrastructure.

  • Work with technical leaders designing and developing scalable and reliable factory components. You will collaborate with multiple AI model teams to understand their requirements to build an efficient infrastructure that improves every teams’ productivity.

  • Define metrics and drive improvements based on user feedback. You will mentor and collaborate throughout the team and with other teams to grow your colleagues and yourself. You will have a history of learning and growing your skills and those around you.

Requirements

  • A history of using your advanced programming skills to build distributed and compute systems, backend services, microservices and cloud technologies.

  • Effective experience working with multi-functional teams, principals and architects, across organizational boundaries.

  • Mentorship, growing teams and team members, and the flexibility to ability to adjust your direction and expectations given the needs of our customers.

  • Deep technical expertise in distributed containerize applications using technologies such as Docker, K8s, Cloud Endpoints, Helm, and Prometheus.

  • Passion for building rich, microservice applications build and test automation pipeline.

  • Excellent interpersonal skills and the ability to lead multi-functional efforts

  • Proven experience debugging and analyzing the performance of distributed microservices or cloud systems.

  • BS or MS in Computer Science, Computer Engineering or related field (or equivalent experience)

  • 8+ years of shown experience developing performant microservice, cloud software and/or tooling roles

Ways to stand out from the crowd:

  • Experience delivering event-driven applications using various services such as Temporal, Kafka, Redis or others and a demonstrable ability to discuss the pros and cons of these choices.

  • A history of building and deploying containers for Microservices, Cloud and On-prem deployments, and their associated CI/CD pipelines

  • Prior experience in working with large scale full stack development

We are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and creative people in the world working for us. If you’re creative and autonomous with a real passion for technology we want to hear from you.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 200,000 USD - 322,000 USD for Level 5.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · WWC 2022

Videos

See all

Related articles

See all