HPC Infrastructure Engineer
TechAffinity Inc
New Haven, United States of America
7 days ago
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
EnglishJob location
New Haven, United States of America
Tech stack
System Configuration
Data Centers
Monitoring of Systems
Linux System Administration
Data Storage Management
High Performance Computing
System Availability
Data Management
Hardware Infrastructure
Job description
We are seeking an experienced HPC Infrastructure Engineer to provide operational and infrastructure support for a High-Performance Computing (HPC) environment. The ideal candidate will have hands-on experience with HPE SGI 8600 systems, HPC cluster administration, enterprise storage, Linux environments, and data center infrastructure. This role requires strong technical expertise along with the ability to coordinate operational activities and support ongoing infrastructure projects., * Provide on-site technical support for the HPE SGI 8600 HPC compute environment and supporting infrastructure.
- Administer, monitor, and maintain HPCM Cluster Management.
- Support the DMF Storage Environment, including system health monitoring, troubleshooting, and upgrade coordination with the storage team.
- Perform maintenance and troubleshooting of CDU (Cooling Distribution Unit) systems.
- Diagnose server hardware issues and coordinate repairs to ensure maximum system availability.
- Work closely with the site administrator to maintain stable and reliable HPC operations.
- Lead weekly operational meetings and coordinate infrastructure-related projects and activities.
- Monitor system performance, identify potential issues, and implement preventive maintenance.
- Maintain documentation related to system configuration, operational procedures, and infrastructure changes.
- Collaborate with cross-functional teams to ensure successful execution of operational initiatives.
- Perform additional infrastructure and operational support responsibilities as needed.
Requirements
- Strong experience supporting High Performance Computing (HPC) environments.
- Hands-on experience with HPE SGI 8600 or similar HPC platforms.
- Experience with HPCM Cluster Management.
- Knowledge of DMF Storage Management and enterprise storage environments.
- Strong Linux system administration skills.
- Experience with enterprise server hardware diagnostics and troubleshooting.
- Familiarity with data center infrastructure, including Cooling Distribution Units (CDUs).
- Experience coordinating infrastructure upgrades, maintenance, and operational support.
- Excellent communication, documentation, and organizational skills.
- Ability to work independently in an onsite production environment., * Experience supporting large-scale enterprise or research computing environments.
- Familiarity with HPC networking and infrastructure best practices.
- Previous experience in a technical lead or infrastructure project coordination role.