HPC Infrastructure Engineer

TechAffinity Inc
New Haven, United States of America
7 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

New Haven, United States of America

Tech stack

System Configuration
Data Centers
Monitoring of Systems
Linux System Administration
Data Storage Management
High Performance Computing
System Availability
Data Management
Hardware Infrastructure

Job description

We are seeking an experienced HPC Infrastructure Engineer to provide operational and infrastructure support for a High-Performance Computing (HPC) environment. The ideal candidate will have hands-on experience with HPE SGI 8600 systems, HPC cluster administration, enterprise storage, Linux environments, and data center infrastructure. This role requires strong technical expertise along with the ability to coordinate operational activities and support ongoing infrastructure projects., * Provide on-site technical support for the HPE SGI 8600 HPC compute environment and supporting infrastructure.

  • Administer, monitor, and maintain HPCM Cluster Management.
  • Support the DMF Storage Environment, including system health monitoring, troubleshooting, and upgrade coordination with the storage team.
  • Perform maintenance and troubleshooting of CDU (Cooling Distribution Unit) systems.
  • Diagnose server hardware issues and coordinate repairs to ensure maximum system availability.
  • Work closely with the site administrator to maintain stable and reliable HPC operations.
  • Lead weekly operational meetings and coordinate infrastructure-related projects and activities.
  • Monitor system performance, identify potential issues, and implement preventive maintenance.
  • Maintain documentation related to system configuration, operational procedures, and infrastructure changes.
  • Collaborate with cross-functional teams to ensure successful execution of operational initiatives.
  • Perform additional infrastructure and operational support responsibilities as needed.

Requirements

  • Strong experience supporting High Performance Computing (HPC) environments.
  • Hands-on experience with HPE SGI 8600 or similar HPC platforms.
  • Experience with HPCM Cluster Management.
  • Knowledge of DMF Storage Management and enterprise storage environments.
  • Strong Linux system administration skills.
  • Experience with enterprise server hardware diagnostics and troubleshooting.
  • Familiarity with data center infrastructure, including Cooling Distribution Units (CDUs).
  • Experience coordinating infrastructure upgrades, maintenance, and operational support.
  • Excellent communication, documentation, and organizational skills.
  • Ability to work independently in an onsite production environment., * Experience supporting large-scale enterprise or research computing environments.
  • Familiarity with HPC networking and infrastructure best practices.
  • Previous experience in a technical lead or infrastructure project coordination role.

Apply for this position