Senior HPC Support Engineer - Ethernet and AI Infrastructure

NVIDIA Corporation
Memphis, TN, United States
1 day ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$108,000.0 - $172,500.0
Working hours
Regular working hours

Tech stack

Link Aggregation (Ethernet) Artificial Intelligence Systems Engineering Bash Shell Border Gateway Protocol Data Centers Cursor (Graphical User Interface Elements) Software Debugging Linux Distributed Systems Ethernet Internet Group Management Protocols
+23 more
Interoperability Python (Programming Language) Routing Network Protocols Open Shortest Path First (OSPF) Overlay Transport Virtualization Ansible Shell Script Transmission Control Protocol (TCP) Tcpdump Wireshark Virtualization Technology YAML AI Infrastructure Computer Network Technologies Deep Learning Containerization Kubernetes Information Technology Routing & Switching Network Server GPT Docker

Job description

We are seeking a highly motivated Senior HPC Support Engineer - Ethernet / AI Infrastructure to be engaging with one of our prestige customers onsite and remote, passionate about data center and networking technologies, to provide comprehensive solutions for sophisticated installations, maintenance, or operations for a broad scope of groundbreaking networking products. You will act as the primary point of contact for this customer. You will spend a minimum of one week per month at the customer site in Memphis, TN, US, supporting technical questions, debugging, and issue resolution. As a member of our NVEX Global Technical Support team, you are a conscientious, proficient communicator who is fundamentally interested in taking ownership in resolving issues, while ensuring that a high level of customer satisfaction is maintained and delivered. Significant part of the role is also to collaborate with Engineering, Marketing, and Support teams regularly on technical issues., * Ability to resolve sophisticated customer concerns and technical issues through meticulous research, reproductions, and solving problems for customers installing our products and supporting systems using Linux Operating Systems (multi-distro), with the focus on NVIDIA Ethernet Switching technologies and our End-to-End Solutions such as NVIDIA Spectrum-X.

  • Responding to customer product support inquiries via telephone, email or conference calls
  • Resolving customer issues during installation, operation, maintenance or product application or interoperability with other vendors
  • Participate in multi-functional team meetings and giving feedback to engineering and marketing regarding product requirements, customer experience, support tools, etc.
  • As a technical resource develop, re-define and document standard methodologies to provide to internal teams(Support/R&D) for support process and improvements

Requirements

  • 5+ years in providing in-depth Customer Support and debugging for hardware and software products.
  • An academic degree from an accredited university or college in Networking, Computer Science/Engineering, or Electrical/IT (or equivalent experience).
  • Shown use of established AI technologies in day-today job responsibilities.
  • Established knowledge of Enterprise platform and systems engineering who understands Linux triage, knowledge about servers and can resolve hardware and/or OS internal issues.
  • Intellectual curiosity, positive attitude, flexibility, analytical ability, self-motivation, and team-oriented including professional-level communication skills, interpersonal skills, with the ability to maintain and lead the overall resolution for any critical issue raised by our customer, under all circumstances.

Profound knowledge and experience (solving) in the following skills:

  • Networking Technology, protocols and routing including TCP, UDP, Ethernet, IP, L2, L3 (ARP, STP, LACP, MLAG, IGMP, PIM, BGP, OSPF), on Enterprise Level.
  • Linux OS including System Administration and Networking (LFCS / RHCSA)
  • Able to debug networking protocols using tools such as TCPDUMP and Wireshark or similar packet generation and analysis tools.
  • Deep understanding of at least two of the following: data centers, servers, distributed systems, virtualization, deep learning frameworks, containers/containerization (i.e. Docker, Kubernetes)
  • Adoption of AI solutions like Cursor, Gemini, ChatGPT, Copilot, Glean, etc. in your daily work routine.

Ways to stand out from the crowd:

Knowledge and working experience with the following:

  • Experience in solving problems in large-scale networking and AI Infrastructure environments with overlay technologies (BGP, OSPF, VXLAN, EVPN), RoCE and QoS Concepts
  • Linux, Networking and NVIDIA AI Infrastructure and Operations Certifications such as CCONP, CCIE, JNCIE-DC/ENT, RHCE, LFCS, NCP-AII/AIO/AIN
  • Shell Scripting (Python, bash, Ansible, yaml, etc…)
  • Effective and comprehensive fixing / debugging methodology

Benefits & conditions

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 108,000 USD - 172,500 USD for Level 3, and 120,000 USD - 207,000 USD for Level 4.

You will also be eligible for equity and benefits.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

1:35 min

Centralizing configuration logic with native YAML block references

Matthieu Vincent Matthieu Vincent · Europe 2026 Virtual

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all