> Markdown version of [/jobs/ext/1180839-high-performance-computing-engineer](https://www.wearedevelopers.com/jobs/ext/1180839-high-performance-computing-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # High Performance Computing Engineer - **Company:** Lyman Products Corporation - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $145,600.0 - **Contract:** Permanent contract - **Skills:** Address Resolution Protocols, Artificial Intelligence, Bash Shell, Command-Line Interface, Computer Clusters, System Configuration, Data Centers, Microprocessors, RAID, Ethernet, Network Interface Controllers, Internet Control Message Protocol, InfiniBand, Subnetting, Job Scheduling, Python (Programming Language), Linux System Administration, Performance Tuning, E2e Testing, Shell Script, TCP/IP, AI Infrastructure, Network Routers, High Performance Computing, Slurm, Hardware Asset Management, Hardware Infrastructure, Terminal Servers - **Published:** July 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=d12b1ae3d1f77a70 ## About the Role Required Skills:· Linux Administration: Strong command-line proficiency and shell scripting (Bash/Python). Experience with performance tuning and troubleshooting.· GPU Infrastructure: Hands-on experience with NVIDIA GPU deployments, driver installation, and end-to-end testing in clustered environments.· High-Speed Networking: Deep working knowledge of InfiniBand (switch config, subnet manager) and Ethernet (TCP/IP, ARP, ICMP).· Hardware: Proficiency in rack & stack, cabling (fiber/copper), and component-level repair (HDDs, RAM, CPUs).· Monitoring: Familiarity with monitoring frameworks and alerting tools.· Soft Skills: Strong documentation skills and ability to coordinate with global teams across time zones. Highly Preferred (Nice to Have):· Experience in HPC, AI/ML, or large-scale supercomputing environments.· Familiarity with job schedulers (e.g., SLURM, PBS).· Experience with data center buildouts or large-scale migrations. Physical Requirements:· Ability to lift and install equipment weighing up to 50lbs.· Comfortable working in data center environments (raised floors, racks, confined spaces).· Willingness to work flexible hours, including nights and weekends for maintenance windows ## Description Job Summary:We are seeking a highly skilled Data Center Operations Engineer to support our cutting-edge HPC and AI infrastructure. This role is heavily focused on Linux system administration, GPU server deployments, and InfiniBand networking. You will be responsible for the hands-on bring-up, maintenance, and troubleshooting of large-scale compute clusters. Key Responsibilities:· Cluster Deployment: Lead the end-to-end bring-up of GPU clusters, including driver installation, system-level configuration, and performance validation.· InfiniBand Management: Perform fabric bring-up, switch configuration, and subnet management for high-performance networks.· Hardware Lifecycle: Install, configure, test, and maintain server hardware (rack & stack, cabling, CPU/GPU, memory, RAID, NICs).· Networking: Configure and troubleshoot routers, switches, and terminal servers (OOB management), including fiber/copper cabling.· Operations & Support: Participate in on-call rotations, conduct daily health checks, and resolve incidents within SLA requirements.· document: Maintain technical runbooks, operational procedures, and system configuration records.· Vendor Management: Coordinate with vendors for hardware delivery, diagnostics, and warranty replacements. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [Debunking the Top 10 Myths about Web 3](https://www.wearedevelopers.com/videos/634-debunking-the-top-10-myths-about-web-3) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Data Science & more: The Lopez dilemma](https://www.wearedevelopers.com/magazine/10-data-science-more-the-lopez-dilemma) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)