Senior Network Reliability Engineer - DGX Cloud

NVIDIA Corporation
Santa Clara, CA, United States
1 day ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$136,000.0 - $224,250.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Border Gateway Protocol Cloud Computing Communications Protocols Complex Networks Wavelength-Division Multiplexing Data Centers Domain Name System (DNS) InfiniBand Networking Hardware Internet Protocol
+21 more
Internet Protocol Security (IP SEC) Multi-protocol Systems Python (Programming Language) Network Troubleshooting Open Shortest Path First (OSPF) Overlay Transport Virtualization Peering Cloud Services Prometheus Shell Script Software Engineering TCP/IP Google Cloud Computer Network Operations Grafana Deep Learning Reliability of Systems Juniper Information Technology Fortinet Oracle Cloud Infrastructure

Job description

  • Engage in 24/7 global shift rotations to provide remote support for network repairs and changes while collaborating across teams and updating customers on status and ticket information.
  • Drive operational improvements in change management and daily operations by following procedures.
  • Manage and operate large scale IP network technologies and infrastructures.
  • Utilize your skills in Peering and Datacenter interconnect technologies: PNI, Transit, Exchange, Passive DWDM, Wave circuits.
  • Monitor and support the network health of on-premises and cloud infrastructures.
  • Collaborate and develop workflow enhancements while documenting best practices.

Requirements

In this role, the Senior Network Operations Engineer will remediate critical alerts within defined SLAs, triage production impacting network incidents, and interact with internal customers on network related issues. They will also be responsible for engaging with external vendors to remediate hardware and software issues, and participate in project related work such as network device upgrades and capacity augmentations. An ideal candidate will possess a wide range of skills, including alert monitoring & resolution in large-scale networks and CSP environments, outstanding troubleshooting skills, understanding of L3 underlay networks, and network protocol knowledge in large multi-vendor infrastructures., * Deep knowledge and experience of TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, and MACsec.

  • 5+ years of experience in network operations.
  • Skilled in network troubleshooting techniques and demonstrating creative problem-solving abilities.
  • Strong track record of alert response within defined SLAs and Incident management.
  • Experience with one or more of the following CSP environments: AWS, Azure, GCP, OCI.
  • Familiarity with Arista, Fortinet and Juniper.
  • Hands-on experience with contributing to tooling and automation for provisioning, monitoring, and managing complex network infrastructures.
  • Bachelor’s degree in Computer Science, related technical field, or equivalent experience.
  • Excellent verbal and written communication skills.

Ways To Stand Out From The Crowd:

  • Solid understanding of Mellanox/Cumulus OS and Infiniband technology.
  • Skilled in Unix/Linux system administration, with the ability to write and understand Python/Shell scripts to improve efficiency in hyperscale environments.
  • Familiarity with leveraging tools such as Netbox/Nautobot, Prometheus, Grafana, Panoptes to monitor and manage a global network. Passionate about innovating and investing in ground breaking technologies.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you. NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these outstanding opportunities in deep learning cloud solutions.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 136,000 USD - 224,250 USD for Level 3, and 168,000 USD - 264,500 USD for Level 4.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

5:28 min

Bitcoin scaling and the transition to peer-to-peer transactions

Jad Wahab · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

1:35 min

Powping protocol for direct peer-to-peer interactions

Glenn Wolfe · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all