Senior Software Engineer, Core Infrastructure Services - DGX Cloud

NVIDIA Ltd.
United States
30 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$168,000.0 - $322,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Border Gateway Protocol Cloud Computing Databases Linux Distributed Systems Domain Name System (DNS) InfiniBand Virtual Private Networks (VPN) Python (Programming Language) Enterprise Messaging Systems
+20 more
Network Architecture Routing OAuth Remote Direct Memory Access Redis Ansible Prometheus Workflow Management Systems Network Routers Load Balancing Grafana Firewalls (Computer Science) Kubernetes Infrastructure Automation Frameworks Iptables Apache Kafka Free and Open-Source Software Amazon Simple Queue Service (SQS) Terraform Microservices

Job description

  • Develop software that enables infrastructure orchestration, self-service workflows, and platform automation.
  • Own integrations with internal and external platforms to automate infrastructure provisioning and lifecycle management.
  • Build observability and security capabilities that improve the reliability and resilience of our infrastructure.
  • Partner with infrastructure and networking teams to deliver production services at scale.
  • Drive operational excellence through automation, monitoring, incident response, and continuous improvement.

Requirements

  • BS or equivalent experience with 8+ years of relevant industry experience.
  • Strong proficiency in Python and Go, with experience building production-quality software.
  • Experience building cloud-native microservices and APIs on Kubernetes using frameworks such as FastAPI, gRPC, or REST.
  • Experience with infrastructure automation (Terraform, Ansible), workflow orchestration (Temporal), and distributed systems using databases, Redis, and messaging platforms (Kafka, NATS, SQS).
  • Experience designing, building, and operating production infrastructure services such as DNS, NTP, AAA (RADIUS/OAuth), and observability platforms.
  • Strong Linux fundamentals with experience in observability (Prometheus, Grafana, OpenTelemetry, gNMI), networking (BGP, switching, routing, load balancing), and security (VPNs, firewalls, iptables/nftables).
  • Excellent problem-solving, communication, and collaboration skills.

Ways to stand out from the crowd:

  • Hands-on experience with network infrastructure including switches, routers, and firewalls.
  • Familiarity with InfiniBand, RDMA, and AI/HPC networking.Experience with NetBox, Nautobot, or similar network source of truth platforms.
  • Contributions to open-source software. Experience with public cloud platforms.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 200,000 USD - 322,000 USD for Level 5.

About the company

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology-and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

NVIDIA is seeking an experienced software engineer to join the Cloud Foundations Automation team. Our team builds and operates the core infrastructure services that power NVIDIA’s DGX Cloud and SuperPod deployments, delivering secure, reliable, and observable platforms at global scale.

What you’ll be doing:

  • Build and operate core infrastructure services that power NVIDIA’s global AI infrastructure.
  • Architect and develop secure, scalable and highly available cloud-native platform services.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:49 min

Adopting OAuth best practices and removing outdated grants

Alexander Schwartz Alexander Schwartz · World Congress 2026 Europe

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all