HPC Engineer
Institute Inc
Sunnyvale, CA, United States
16 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.thejobnetwork.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Microsoft Azure
Bash Shell
Cloud Computing
Computer Clusters
Computer Engineering
Linux
DevOps
Disaster Recovery
Monitoring of Systems
Python (Programming Language)
Linux System Administration
+11 more
Reliability Engineering
Prometheus
Software Engineering
Datadog
Scripting
High Performance Computing
Grafana
Kubernetes
Information Technology
Slurm
Hardware Infrastructure
Job description
This role provides operational coverage during Abu Dhabi overnight hours and serves as a primary point of contact for infrastructure monitoring, incident triage, researcher support, and production operations. Responsibilities * Monitor health, performance, and availability of large-scale GPU clusters.
- Respond to incidents and perform first-level triage.
- Support researchers and troubleshoot job failures.
- Execute operational runbooks and recovery procedures.
- Validate cluster deployments, upgrades, and maintenance activities.
- Track infrastructure utilization and operational metrics.
- Develop automation and monitoring tools.
Requirements
- Contribute to documentation and reporting. Education Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, Information Technology, Electrical Engineering, Mathematics, Physics, or related disciplines.
Experience * 2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC, or infrastructure operations.
- Strong Linux troubleshooting skills.
- Experience with scripting using Python or Bash. Preferred Qualifications * Slurm.
- GPU infrastructure.
- AWS, Azure, or GCP.
- Grafana, Prometheus, Datadog, or similar tools.
- Containers and Kubernetes.
Benefits & conditions
- Research computing environments. Salary Range The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs. Benefits Include *Comprehensive medical, dental, and vision benefits *Bonus *401K Plan *Generous paid time off, sick leave and holidays *Paid Parental Leave *Employee Assistance Program *Life insurance and disability
About the company
The Institute for Foundation Models (IFM), operates some of the world’s largest AI supercomputing environments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.thejobnetwork.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
IK
Igor Khokhriakov
about 2 months ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
LM
Luis Minvielle
Top 6 Hackathons for Developers in 2023
about 3 years ago
EM
Eli McGarvie
DevOps Engineer Salary [2023]
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
over 2 years ago