High Performance Computing Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+16 more
Job description
Job Summary:We are seeking a highly skilled Data Center Operations Engineer to support our cutting-edge HPC and AI infrastructure. This role is heavily focused on Linux system administration, GPU server deployments, and InfiniBand networking. You will be responsible for the hands-on bring-up, maintenance, and troubleshooting of large-scale compute clusters.
Key Responsibilities:· Cluster Deployment: Lead the end-to-end bring-up of GPU clusters, including driver installation, system-level configuration, and performance validation.· InfiniBand Management:
Perform fabric bring-up, switch configuration, and subnet management for high-performance networks.· Hardware Lifecycle: Install, configure, test, and maintain server hardware (rack & stack, cabling, CPU/GPU, memory, RAID, NICs).·
Networking: Configure and troubleshoot routers, switches, and terminal servers (OOB management), including fiber/copper cabling.·
Operations & Support: Participate in on-call rotations, conduct daily health checks, and resolve incidents within SLA requirements.·
document: Maintain technical runbooks, operational procedures, and system configuration records.· Vendor Management: Coordinate with vendors for hardware delivery, diagnostics, and warranty replacements.
Requirements
Required Skills:· Linux Administration: Strong command-line proficiency and shell scripting (Bash/Python). Experience with performance tuning and troubleshooting.· GPU Infrastructure: Hands-on experience with NVIDIA GPU deployments, driver installation, and end-to-end testing in clustered environments.·
High-Speed Networking: Deep working knowledge of InfiniBand (switch config, subnet manager) and Ethernet (TCP/IP, ARP, ICMP).· Hardware: Proficiency in rack & stack, cabling (fiber/copper), and component-level repair (HDDs, RAM, CPUs).·
Monitoring: Familiarity with monitoring frameworks and alerting tools.· Soft Skills: Strong documentation skills and ability to coordinate with global teams across time zones.
Highly Preferred (Nice to Have):· Experience in HPC, AI/ML, or large-scale supercomputing environments.· Familiarity with job schedulers (e.g., SLURM, PBS).· Experience with data center buildouts or large-scale migrations.
Physical Requirements:· Ability to lift and install equipment weighing up to 50lbs.· Comfortable working in data center environments (raised floors, racks, confined spaces).· Willingness to work flexible hours, including nights and weekends for maintenance windows
Benefits & conditions
$70 an hour - Contract
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Top 6 Hackathons for Developers in 2023
Data Science & more: The Lopez dilemma
Dev Digest 121 - AI goes offline
Dev Digest 120 - Apple and peers