GPU Server & Systems Operations Engineer

Nexo Global Inc
Irvine, CA, United States
8 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$65,000.0 - $90,000.0
Working hours
Regular working hours
Job source

Tech stack

Intelligent Platform Management Interface BIOS Computer Clusters Nvidia CUDA Linux RAID Hard Disk Drives Firmware Monitoring of Systems Networking Basics Prometheus Server Administration
+8 more
Shell Script Zabbix Graphics Processing Unit (GPU) Grafana Containerization Kubernetes Slurm Hardware Infrastructure

Job description

  • Conduct acceptance, rack installation, cabling, initialization and asset registration upon arrival of GPU servers.
  • Troubleshoot hardware faults involving GPUs, memory, hard drives, power supplies, network adapters, optical modules and other components.
  • Administer BMC, IPMI, BIOS, RAID, firmware and fundamental Linux operating systems.
  • Install and maintain foundational environments including NVIDIA drivers, CUDA and DCGM.
  • Collaborate with the Shenzhen team to resolve failures related to Slurm, Kubernetes, storage and GPU clusters.
  • Manage server monitoring alerts, routine inspections, spare parts inventory and vendor RMA processes.
  • Assist the datacenter with power supply, temperature and liquid cooling alerts; compile operation & maintenance SOPs and conduct fault RCA.
  • Participate in emergency on-call rotation and on-site incident resolution.

Requirements

  • Minimum 4 years of working experience in server, IDC or Linux systems operations.
  • Familiar with GPU server hardware, BMC/IPMI, BIOS, RAID and firmware upgrades.
  • Proficient in Linux, Shell scripting and basic networking knowledge.
  • Working knowledge of NVIDIA GPU, CUDA, DCGM, NVLink/NVSwitch.
  • Familiar with monitoring tools such as Prometheus, Grafana and Zabbix.
  • Experience with Slurm, Kubernetes or container platforms is preferred.
  • Hands-on experience with B300/GB200/H100/H200 or liquid-cooled servers is preferred.
  • English communication skills for interactions with datacenter staff and vendors; must possess valid legal work authorization in the United States.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role β€” technically off-topic, practically not.

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 Β· LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard Β· WWC 2025

1:04 min

Core components driving the Nvidia GPU operator

Kevin Klues Kevin Klues

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 Β· Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani Β· Europe 2026 Virtual

8:22 min

Simulating a Linux terminal and running Spring Boot

Jakov Semenski Β· LIVE

Videos

See all

Related articles

See all