HPC SYSTEMS ENGINEER

University of Washington
Seattle, WA, United States
12 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$97,080.0 - $157,764.0
Working hours
Shift work
Job source

Tech stack

Link Aggregation (Ethernet) Systems Engineering Backup Devices Bash Shell CentOS Computer Programming System Configuration Data Centers Database Queries Linux Web Development Django Web Framework
+38 more
Elasticsearch Ethernet Firmware Systems Theories General Parallel File Systems Monitoring of Systems IBM Storage Issue Tracking Systems InfiniBand Networking Hardware Python (Programming Language) PostgreSQL Linux System Administration MariaDB MongoDB MySQL NoSQL Red Hat Enterprise Linux Redis Ansible Prometheus Virtual Local Area Networks Ceph (Software) Scripting Saltstack Grafana Git Containerization Infrastructure Automation Frameworks Information Technology Deployment Automation Non-relational Database Slurm Lxc Puppet Software Version Control Docker Golang

Job description

Reporting to Director, the HPC Systems Engineer is responsible for supporting the HPC efforts for the research computing group and related research computing technologies. This position also provids expertise and support to other endeavors when needed. This position requires a team-oriented professional, experienced in managing large IT systems for research using automated and software-defined approaches. This position regularly interfaces with other UWIT teams as well as research customers across campus. This position requires participation in a 24x7 on-call rotation. This is an essential position is required to work remotely when the University suspends operations.

Key Duties: 25% Systems Engineering and Development: Design, develop, adapt, integrate, manage, optimize, or deploy monitoring for new or existing systems, software, or automation to continually improve the performance, reliability, recoverability, manageability, and usability of research cyberinfrastructure.

20% Systems Administration: Perform system management functions such as account administration, service administration, software installation and configuration, firmware updates, software updates, backing up and restoring data, improving system and service monitoring, or planning and execution of hardware replacements.

15% Operational Monitoring and Troubleshooting: Monitor HPC and related systems health and performance, and troubleshoot and resolve issues that impact performance, health, or reliability.

20% Technical Support: Monitor support ticket and alerting queue; triage, respond to, and resolve tickets and on-call pages as appropriate.

10% Systems Research, Evaluation and Architecting: Research and evaluate new and future technologies. Plan and architect new systems and the evolution and lifecycle of existing systems and clusters. Suggest or advise on new technologies to the Director of Research Computing Operations and other members of the Research Computing Operations team. 5% Training, Mentoring and Documentation: Effectively share experience and knowledge with other team members. 5% Other duties as assigned Responsibilities

Requirements

To be considered for this opportunity your application must demonstrate you meet both the minimum qualifications and additional qualifications listed below. Equivalent education and/or experience may substitute for minimum qualifications except when there are legal requirements, such as a license, certification, and/or registration., Bachelors degree in computer science, information technology, scientific, engineering, or related field or experience Minimum 4 years’ experience in Linux system administration experience or substantial experience working with Linux. Basic knowledge of networking hardware, software, protocols, and concepts. Demonstrated ability to work with minimal supervision, both independently and as part of a team. Demonstrated excellent written/oral communication skills, technical documentation skills, user liaison skills, and personal interaction abilities. Applicants who do not meet these qualifications WILL NOT be forwarded to the Hiring Manager.

Preferred Qualifications Knowledge of Containerization platforms (e.g., Docker, Apptainer, LXC, Podman). Web development skills such as Python and Django. Security compliance experience such as NIST 800-171, CMMC L2, and FAR to protect data such as HIPAA, CUI, Export Controlled, etc. Progressively responsible experience as an engineer, architect, or role with comparable technical responsibilities in a large Linux HPC environment. Extensive experience with administration of Linux operating systems in a production environment, including experience with Red Hat Enterprise Linux or derivatives such as CentOS or Rocky Linux. Familiarity with SLURM or other HPC scheduler (PBS Pro, PBS/Torque, SGE/UGE, LSF, etc). Experience designing, configuring, and troubleshooting networks using both Ethernet and high-performance interconnects such as Infiniband (e.g. experience with VLAN, MLAG, and LACP configurations). Proficiency in programming/scripting languages in the context of systems engineering or administration, preferably including Bash, Golang, or Python. Experience in the configuration and use of mass deployment tools such as Warewulf, MAAS, Foreman, xCAT, Cobbler, or similar. Proficiency with the use of Git for source control in collaboration with a team with multiple contributors. Ability to administer and troubleshoot large high-performance parallel filesystems such as IBM Storage Scale (GPFS), Lustre, BeeGFS, Ceph. Experience with the use of configuration management tools such as Ansible, SaltStack, or Puppet. Experience with monitoring, analysis and visualization of metrics, telemetry, and logs (e.g. Grafana, Prometheus, splunc elasticsearch, rsyslog, etc). Experience in a data center environment (e.g., racking equipment, running cables, labeling, asset tracking). Experience with AI training and inference. Experience with timeseries, relational, and non-relational database management or database query construction (e.g. influx, promql, redis, mongodb, postgres, mysql, nosql, mariadb, etc). Scientific background, research experience, and/or experience in a University setting.

Benefits & conditions

Requires monitoring of e-mail and trouble ticket system for questions needing immediate response during business hours. On-call responsibilities for after-hours system outages. Server management will include both production (24x7x365) and development systems. Open office environment. Hybrid - expect to be in the office a minimum of two days per week. This position requires participation in a 24x7 on-call rotation. This is an essential position is required to work remotely when the University suspends operations.

About the Team UWIT is the University of Washington’s central IT organization, supporting all three campuses, UW medical centers, and global research. We partner across the University to power teaching, learning, innovation, and discovery. We are here to serve UW. The Research Computing Operations group within the UWIT, designs, implements, and operates Linux-based High-Performance Computing (HPC) clusters and storage systems to offer UW researchers cost-effective yet powerful high-performance computing options that unlock new possibilities for their research.

Compensation, Benefits and Position Details

Pay Range Minimum: $97,080.00 annual Pay Range Maximum: $157,764.00 annual Other Compensation, For information about benefits for this position, visit ;br>Shift: First Shift (United States of America) Temporary or Regular? This is a regular position FTE (Full-Time Equivalent): 100.00% Union/Bargaining Unit: Not Applicable

About the company

Working at the University of Washington provides a unique opportunity to change lives - on our campuses, in our state and around the world.

UW employees bring their boundless energy, creative problem-solving skills and dedication to building stronger minds and a healthier world. In return, they enjoy outstanding benefits, opportunities for professional growth and the chance to work in an environment known for its diversity, intellectual excitement, artistic pursuits and natural beauty.

Our Commitment

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:55 min

Enterprise container migration and historical technology evolution

Federico Fregosi · World Congress 2022

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all