GPU Server & Systems Operations Engineer
Nexo Global Inc
Irvine, CA, United States
8 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$65,000.0 - $90,000.0
Working hours
Regular working hours
Job source
Tech stack
Intelligent Platform Management Interface
BIOS
Computer Clusters
Nvidia CUDA
Linux
RAID
Hard Disk Drives
Firmware
Monitoring of Systems
Networking Basics
Prometheus
Server Administration
+8 more
Shell Script
Zabbix
Graphics Processing Unit (GPU)
Grafana
Containerization
Kubernetes
Slurm
Hardware Infrastructure
Job description
- Conduct acceptance, rack installation, cabling, initialization and asset registration upon arrival of GPU servers.
- Troubleshoot hardware faults involving GPUs, memory, hard drives, power supplies, network adapters, optical modules and other components.
- Administer BMC, IPMI, BIOS, RAID, firmware and fundamental Linux operating systems.
- Install and maintain foundational environments including NVIDIA drivers, CUDA and DCGM.
- Collaborate with the Shenzhen team to resolve failures related to Slurm, Kubernetes, storage and GPU clusters.
- Manage server monitoring alerts, routine inspections, spare parts inventory and vendor RMA processes.
- Assist the datacenter with power supply, temperature and liquid cooling alerts; compile operation & maintenance SOPs and conduct fault RCA.
- Participate in emergency on-call rotation and on-site incident resolution.
Requirements
- Minimum 4 years of working experience in server, IDC or Linux systems operations.
- Familiar with GPU server hardware, BMC/IPMI, BIOS, RAID and firmware upgrades.
- Proficient in Linux, Shell scripting and basic networking knowledge.
- Working knowledge of NVIDIA GPU, CUDA, DCGM, NVLink/NVSwitch.
- Familiar with monitoring tools such as Prometheus, Grafana and Zabbix.
- Experience with Slurm, Kubernetes or container platforms is preferred.
- Hands-on experience with B300/GB200/H100/H200 or liquid-cooled servers is preferred.
- English communication skills for interactions with datacenter staff and vendors; must possess valid legal work authorization in the United States.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role β technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
MH
Michael Hunger
Everything a Developer Needs to Know About MCP with Neo4j
about 1 year ago
CS
Christina Schaireiter
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
2 months ago
LM
Luis Minvielle
A Guide to Green Tech and Green IT Careers
over 2 years ago
DA
Dr. Andy R. Terrel - NVIDIA
Whatβs the latest in NVIDIA CUDA Python
over 1 year ago