HPC Engineer

Institute Inc
Sunnyvale, CA, United States
16 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Computer Clusters Computer Engineering Linux DevOps Disaster Recovery Monitoring of Systems Python (Programming Language) Linux System Administration
+11 more
Reliability Engineering Prometheus Software Engineering Datadog Scripting High Performance Computing Grafana Kubernetes Information Technology Slurm Hardware Infrastructure

Job description

This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations. Responsibilities * Monitor health, performance, and availability of large-scale GPU clusters.

  • Respond to incidents and perform first-level triage.
  • Support researchers and troubleshoot job failures.
  • Execute operational runbooks and recovery procedures.
  • Validate cluster deployments, upgrades, and maintenance activities.
  • Track infrastructure utilization and operational metrics.
  • Develop automation and monitoring tools.

Requirements

  • Contribute to documentation and reporting. Education Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, Information Technology, Electrical Engineering, Mathematics, Physics, or related disciplines.

Experience * 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.

  • Strong Linux troubleshooting skills.
  • Experience with scripting using Python or Bash. Preferred Qualifications * Slurm.
  • GPU infrastructure.
  • AWS, Azure, or GCP.
  • Grafana, Prometheus, Datadog, or similar tools.
  • Containers and Kubernetes.

Benefits & conditions

  • Research computing environments. Salary Range The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs. Benefits Include *Comprehensive medical, dental, and vision benefits *Bonus *401K Plan *Generous paid time off, sick leave and holidays *Paid Parental Leave *Employee Assistance Program *Life insurance and disability

About the company

The Institute for Foundation Models (IFM), operates some of the world’s largest AI supercomputing environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.thejobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

Videos

See all

Related articles

See all