HPC Infrastructure Platform Engineer

Cadre5, LLC.
Knoxville, TN, United States
27 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Clean Code Principles Bash Shell Configuration Management Code Review Dynamic Host Configuration Protocol Linux Domain Name System (DNS) Github Python (Programming Language) Network Protocols OpenLDAP Ansible
+10 more
Management of Software Versions Virtual Environment Gitlab Kubernetes Infrastructure Automation Frameworks Information Technology Puppet Golang Vmware Programming Languages

Job description

Linux Administration:

  • Deploy, configure and manage HPC-scale services in a Linux environment, primarily RedHat and Rocky
  • Perform regular patches, updates and backups
  • Monitor systems using tools like Nagios and Grafana
  • Respond to and assist in troubleshooting issues

Kubernetes Administration and Automation:

  • Build and maintain foundational internal platforms and tools to enable the HPC Infrastructure team to reliably deploy, monitor and scale applications
  • Design standardized and automated workflow patterns, build and maintain CI/CD pipelines
  • Offer self-service, excellent documentation and assistance to HPC Infrastructure group members for efficient consumption of platform services
  • Develop, maintain and review high-quality code for internal tools using programming languages such as Python, Golang, or Rust
  • Define policies and procedures for automation and configuration management for the team and organization as a whole

Project Management and Leadership:

  • Lead small Infrastructure projects through the project lifecycle
  • Mentor and train junior staff, creating training documentation, holding knowledge sharing sessions, and fostering skill growth throughout the team
  • Propose and implement improvements to existing Infrastructure systems as well as new systems, processes and procedures

Requirements

  • Bachelor’s degree in computer science or closely related field and a minimum of 5 years of experience in Linux systems and Kubernetes platform administration, or a master’s degree and a minimum of 4 years of experience in Linux systems and Kubernetes platform administration
  • An equivalent combination of education and experience will be considered
  • The ability to obtain and maintain a Department of Energy “Q” clearance is required. This requires US Citizenship.

Preferred Qualifications:

  • Excellent interpersonal/communication skills and the ability to work within a team
  • Strong experience designing, building and maintaining Kubernetes platform tools
  • Strong working knowledge of Linux system fundamentals and common network protocols
  • Programming and scripting skills in common languages such as Python and bash
  • Understanding of versioning and code review tools like GitHub and GitLab
  • Experience implementing and supporting highly available systems and services
  • Experience with configuration management tools such as Puppet or Ansible
  • Experience deploying and maintaining virtual environments using VMWare
  • Experience deploying, maintaining and troubleshooting a variety of infrastructure services such as OpenLDAP, DNS, DHCP, etc.
  • Ability to plan, prioritize and complete assigned projects with minimal supervision

About the company

Founded in 1999 in the beautiful Smoky Mountains of East Tennessee, Cadre5 provides innovative technical solutions to our customers locally and nationally. Our Cadre5 Lab Partners division has partnered with The High-Performance Computing Systems Section within the National Center for Computational Sciences (NCCS) at Oak Ridge National Laboratory (ORNL) to recruit an HPC Infrastructure Platform Engineer to join the HPC Infrastructure group. The preferred candidate will possess commensurate knowledge, skills and abilities in addition to relevant education, certifications, experience and demonstrated ability to work as a member of a team. NCCS Provides state-of-the-art computational and data science infrastructure coupled with dedicated technical and scientific professionals tackling large-scale problems across a broad range of scientific domains for accelerating scientific discovery and engineering advances. NCCS hosts the Oak Ridge Leadership Computing Facility (OLCF), one of the Department of Energy’s (DOE) National User Facilities which operates Frontier, the nation’s first exascale supercomputer. ORNL delivers scientific discoveries and technical breakthroughs needed to realize solutions in energy and national security and provides economic benefit to the nation. This premier research institution located near Knoxville in Oak Ridge, TN, addresses national needs through impactful research and world-leading research centers. This is a full-time, permanent position that can telecommute. The ideal candidate will be willing to relocate to the Knoxville, TN/Oak Ridge, TN area and work on-site as needed.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · World Congress 2024

Videos

See all

Related articles

See all