> Markdown version of [/jobs/ext/2312739-sr-hpc-systems-architect-linux](https://www.wearedevelopers.com/jobs/ext/2312739-sr-hpc-systems-architect-linux). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. HPC Systems Architect Linux - **Company:** KLATencor Corporation - **Location:** Ann Arbor, MI, United States - **Experience:** Expert - **Salary:** $129,600.0 - $220,300.0 - **Contract:** Permanent contract - **Skills:** Apache HTTP Server, BIOS, Ubuntu (Operating System), Configuration Management, Dynamic Host Configuration Protocol, Linux, DevOps, File Systems, Domain Name System (DNS), General Parallel File Systems, Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Python (Programming Language), Lightweight Directory Access Protocols (LDAP), Networking Basics, Nginx, Red Hat Enterprise Linux, Ansible, Prometheus, TCP/IP, Scripting, Graphics Processing Unit (GPU), High Performance Computing, SUSE Linux, Grafana, Git, Puppet, Network Server, Docker - **Published:** August 30, 2026 - **Apply:** https://jobs.mitalent.org/job-seeker/job-details/JobCode/405167761 ## About the Role Requires minimum of 8 years of related experience with a Bachelor's degree; or 6 years and a Master's degree; or a PhD with 3 years experience; or equivalent experience. ## Description This role provides senior technical leadership for the architecture, deployment, and longterm scalability of largescale HPC storage and compute platforms! It owns systems endtoend-from early architectural definition through full production-partnering across engineering, manufacturing, and strategic vendors to deliver highly available, highperformance infrastructure at scale. The scope emphasizes deep technical ownership, architectural decisionmaking, and solving sophisticated infrastructure challenges in live production environments! This work directly develops critically important HPC platforms built for adaptability, scale, and operational excellence, driving realworld impact across core products and technologies.Job Duties, but not limited to: Lead the design, implementation, and ongoing support of highperformance compute (HPC) clusters, taking accountability for system performance, reliability, and scalability Serve as a technical authority for HPC storage, with deep handson expertise in parallel file systems such as Lustre, GPFS, and BeeGFS Apply sophisticated systems knowledge across CPU and GPU architectures, highbandwidth interconnects, and robust storage subsystems to deliver balanced, highperformance solutions Lead the creation of hardware BOMs for HPC clusters, working directly with vendors and coordinating hardware release activities Design, configure, and optimize Linux operating systems for HPC environments. Translate project specifications and performance requirements into subsystem and systemlevel designs, driving execution while meeting technical and schedule commitments Support the design, release, and transition of new systems to manufacturing and customers, providing highquality golden images, procedures, scripts, and documentation Lead EOL part requalification activities to ensure longterm system viability and supportabilityQualifications, but not limited to: Proven experience with HPC systems and Linux platform. Strong, distroagnostic Linux experience (Rocky, RHEL, SuSE, Ubuntu) Strong scripting skills in Shell and Python Strong understanding of HPC hardware platforms (servers, GPUs, networking, storage, BIOS/BMC) Advanced Linux systems knowledge (PXE/netboot, systemd, HA concepts) Solid networking fundamentals (TCP/IP, DNS, DHCP, LDAP, HTTP) Experience with configuration management and automation (Salt, Ansible, Puppet, Chef, etc.) Interest in HPC storagePreferred Qualifications: Strong DevOps and automation mentality (CI/CD pipelines, Git, infrastructure as code) Experience with containers for HPC (Singularity, Docker) Monitoring and observability experience (Prometheus, Grafana) Familiarity with Apache/Nginx and supporting infrastructure services ## Related Videos - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [The best of two worlds - Bringing enterprise-grade Linux to the vehicle](https://www.wearedevelopers.com/videos/67-the-best-of-two-worlds-bringing-enterprise-grade-linux-to-the-vehicle) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)