High Performance Computing Engineer

Lyman Products Corporation
San Jose, CA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$145,600.0
Working hours
Shift work
Job source

Tech stack

Address Resolution Protocols Artificial Intelligence Bash Shell Command-Line Interface Computer Clusters System Configuration Data Centers Microprocessors RAID Ethernet Network Interface Controllers Internet Control Message Protocol
+16 more
InfiniBand Subnetting Job Scheduling Python (Programming Language) Linux System Administration Performance Tuning E2e Testing Shell Script TCP/IP AI Infrastructure Network Routers High Performance Computing Slurm Hardware Asset Management Hardware Infrastructure Terminal Servers

Job description

Job Summary:We are seeking a highly skilled Data Center Operations Engineer to support our cutting-edge HPC and AI infrastructure. This role is heavily focused on Linux system administration, GPU server deployments, and InfiniBand networking. You will be responsible for the hands-on bring-up, maintenance, and troubleshooting of large-scale compute clusters.

Key Responsibilities:· Cluster Deployment: Lead the end-to-end bring-up of GPU clusters, including driver installation, system-level configuration, and performance validation.· InfiniBand Management:

Perform fabric bring-up, switch configuration, and subnet management for high-performance networks.· Hardware Lifecycle: Install, configure, test, and maintain server hardware (rack & stack, cabling, CPU/GPU, memory, RAID, NICs).·

Networking: Configure and troubleshoot routers, switches, and terminal servers (OOB management), including fiber/copper cabling.·

Operations & Support: Participate in on-call rotations, conduct daily health checks, and resolve incidents within SLA requirements.·

document: Maintain technical runbooks, operational procedures, and system configuration records.· Vendor Management: Coordinate with vendors for hardware delivery, diagnostics, and warranty replacements.

Requirements

Required Skills:· Linux Administration: Strong command-line proficiency and shell scripting (Bash/Python). Experience with performance tuning and troubleshooting.· GPU Infrastructure: Hands-on experience with NVIDIA GPU deployments, driver installation, and end-to-end testing in clustered environments.·

High-Speed Networking: Deep working knowledge of InfiniBand (switch config, subnet manager) and Ethernet (TCP/IP, ARP, ICMP).· Hardware: Proficiency in rack & stack, cabling (fiber/copper), and component-level repair (HDDs, RAM, CPUs).·

Monitoring: Familiarity with monitoring frameworks and alerting tools.· Soft Skills: Strong documentation skills and ability to coordinate with global teams across time zones.

Highly Preferred (Nice to Have):· Experience in HPC, AI/ML, or large-scale supercomputing environments.· Familiarity with job schedulers (e.g., SLURM, PBS).· Experience with data center buildouts or large-scale migrations.

Physical Requirements:· Ability to lift and install equipment weighing up to 50lbs.· Comfortable working in data center environments (raised floors, racks, confined spaces).· Willingness to work flexible hours, including nights and weekends for maintenance windows

Benefits & conditions

$70 an hour - Contract

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · WWC 2022

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · WWC Europe 2026

2:51 min

Designing hardware infrastructure and networking for distributed compute

Anshul Jindal Anshul Jindal +1 · WWC 2025

Videos

See all

Related articles

See all