> Markdown version of [/jobs/ext/3313102-hpc-systems-engineer](https://www.wearedevelopers.com/jobs/ext/3313102-hpc-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC SYSTEMS ENGINEER - **Company:** University of Washington - **Location:** Seattle, WA, United States (Remote available) - **Experience:** Experienced - **Salary:** $97,080.0 - $157,764.0 - **Contract:** Temporary contract - **Skills:** Link Aggregation (Ethernet), Systems Engineering, Backup Devices, Bash Shell, CentOS, Computer Programming, System Configuration, Data Centers, Database Queries, Linux, Web Development, Django Web Framework, Elasticsearch, Ethernet, Firmware, Systems Theories, General Parallel File Systems, Monitoring of Systems, IBM Storage, Issue Tracking Systems, InfiniBand, Networking Hardware, Python (Programming Language), PostgreSQL, Linux System Administration, MariaDB, MongoDB, MySQL, NoSQL, Red Hat Enterprise Linux, Redis, Ansible, Prometheus, Virtual Local Area Networks, Ceph (Software), Scripting, Saltstack, Grafana, Git, Containerization, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Non-relational Database, Slurm, Lxc, Puppet, Software Version Control, Docker, Golang - **Published:** September 20, 2026 - **Apply:** https://www.dice.com/job-detail/d43cd7ad-90dd-4b8d-8b6d-0f96585cb805 ## About the Role To be considered for this opportunity your application must demonstrate you meet both the minimum qualifications and additional qualifications listed below. Equivalent education and/or experience may substitute for minimum qualifications except when there are legal requirements, such as a license, certification, and/or registration., Bachelors degree in computer science, information technology, scientific, engineering, or related field or experience Minimum 4 years' experience in Linux system administration experience or substantial experience working with Linux. Basic knowledge of networking hardware, software, protocols, and concepts. Demonstrated ability to work with minimal supervision, both independently and as part of a team. Demonstrated excellent written/oral communication skills, technical documentation skills, user liaison skills, and personal interaction abilities. Applicants who do not meet these qualifications WILL NOT be forwarded to the Hiring Manager. Preferred Qualifications Knowledge of Containerization platforms (e.g., Docker, Apptainer, LXC, Podman). Web development skills such as Python and Django. Security compliance experience such as NIST 800-171, CMMC L2, and FAR to protect data such as HIPAA, CUI, Export Controlled, etc. Progressively responsible experience as an engineer, architect, or role with comparable technical responsibilities in a large Linux HPC environment. Extensive experience with administration of Linux operating systems in a production environment, including experience with Red Hat Enterprise Linux or derivatives such as CentOS or Rocky Linux. Familiarity with SLURM or other HPC scheduler (PBS Pro, PBS/Torque, SGE/UGE, LSF, etc). Experience designing, configuring, and troubleshooting networks using both Ethernet and high-performance interconnects such as Infiniband (e.g. experience with VLAN, MLAG, and LACP configurations). Proficiency in programming/scripting languages in the context of systems engineering or administration, preferably including Bash, Golang, or Python. Experience in the configuration and use of mass deployment tools such as Warewulf, MAAS, Foreman, xCAT, Cobbler, or similar. Proficiency with the use of Git for source control in collaboration with a team with multiple contributors. Ability to administer and troubleshoot large high-performance parallel filesystems such as IBM Storage Scale (GPFS), Lustre, BeeGFS, Ceph. Experience with the use of configuration management tools such as Ansible, SaltStack, or Puppet. Experience with monitoring, analysis and visualization of metrics, telemetry, and logs (e.g. Grafana, Prometheus, splunc elasticsearch, rsyslog, etc). Experience in a data center environment (e.g., racking equipment, running cables, labeling, asset tracking). Experience with AI training and inference. Experience with timeseries, relational, and non-relational database management or database query construction (e.g. influx, promql, redis, mongodb, postgres, mysql, nosql, mariadb, etc). Scientific background, research experience, and/or experience in a University setting. ## Description Reporting to Director, the HPC Systems Engineer is responsible for supporting the HPC efforts for the research computing group and related research computing technologies. This position also provids expertise and support to other endeavors when needed. This position requires a team-oriented professional, experienced in managing large IT systems for research using automated and software-defined approaches. This position regularly interfaces with other UWIT teams as well as research customers across campus. This position requires participation in a 24x7 on-call rotation. This is an essential position is required to work remotely when the University suspends operations. Key Duties: 25% Systems Engineering and Development: Design, develop, adapt, integrate, manage, optimize, or deploy monitoring for new or existing systems, software, or automation to continually improve the performance, reliability, recoverability, manageability, and usability of research cyberinfrastructure. 20% Systems Administration: Perform system management functions such as account administration, service administration, software installation and configuration, firmware updates, software updates, backing up and restoring data, improving system and service monitoring, or planning and execution of hardware replacements. 15% Operational Monitoring and Troubleshooting: Monitor HPC and related systems health and performance, and troubleshoot and resolve issues that impact performance, health, or reliability. 20% Technical Support: Monitor support ticket and alerting queue; triage, respond to, and resolve tickets and on-call pages as appropriate. 10% Systems Research, Evaluation and Architecting: Research and evaluate new and future technologies. Plan and architect new systems and the evolution and lifecycle of existing systems and clusters. Suggest or advise on new technologies to the Director of Research Computing Operations and other members of the Research Computing Operations team. 5% Training, Mentoring and Documentation: Effectively share experience and knowledge with other team members. 5% Other duties as assigned Responsibilities ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [Stop Committing Your Secrets - GIt Hooks To The Rescue!](https://www.wearedevelopers.com/videos/573-stop-committing-your-secrets-git-hooks-to-the-rescue) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)