Senior HPC DevOPS Engineer | TS/SCI w/MD poly required

Powerschool Group LLC
College Park, MD, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$96,400.0 - $144,600.0
Working hours
Regular working hours

Tech stack

Agile Methodology Artificial Intelligence Systems Engineering Confluence Automation of Tests Cyber Security Linux DevOps IBM Hardware Management Console Linux System Administration Node.Js Role-Based Access Control
+11 more
Release Management Ansible Runbook Management of Software Versions Virtualization Technology Data Logging Software Troubleshooting Git Selinux Information Technology Bare Metal

Job description

We are seeking a fully cleared Senior HPC DevOps Engineer to own the operations and automation lifecycle for an existing HPC/AI compute cluster (Linux). You will work closely with team members, as well as directly with our Maryland-based customer, in a fast-paced environment. In this role you will codify repeatable operations in Ansible and drive execution through an enterprise automation controller to enforce desired state, detect drift, accelerate node onboarding, and streamline incident response via runbook automation integrated with monitoring and ITSM. Key Responsibilities may include: Automation ownership: Job templates, inventories, credentials, RBAC, execution environments, promotion across environments. Desired-state and drift detection: Enforce desired state across cluster services via code-driven configuration; implement drift detection and alert on deviations; reconcile runtime state vs configured state. Compute node onboarding (Bare-metal/VM): Build and maintain an automated node bootstrap workflow that installs/configures the OS, applies security and performance baselines, enrolls nodes into the scheduler and shared storage ecosystem, validates hardware and service readiness (CPU, network, accelerator, storage mounts), and reports pass/fail results. Patch & vulnerability response: Implement rolling maintenance and patch automation to meet defined vulnerability response SLAs. Maintain version-controlled container build definitions and integrate image scanning into the build/release lifecycle. Logging & observability: Ensure automation and operational workflows emit auditable logs to centralized analytics and integrate with metrics/alerting to enable reliable incident response, proactive detection, and safe auto-remediation. Incident/problem management: Automate responses to common incidents (hung nodes, storage performance alarms, image vulnerabilities, hardware failures) leveraging out-of-band hardware management interfaces and standardized runbooks. Docs-as-code: Keep runbooks and operational documentation versioned alongside automation and publish operator guidance to the orgs documentation platform., Job Summary: The DevOps Engineer is a technical operational engineering role responsible for improving DevOps, Release Management, Test Automation, and platform operations throug…

  • 4 days ago +

Requirements

12+ years of experience and a BS in computer science, IT, or related technical field, MS and 10 years of experience, or a Ph.D. with 8 years of experience. Four years of additional experience is required in lieu of a Bachelors’ degree for a total of 16 years of experience. 7+ years in Linux systems / SRE / DevOps, including production cluster operations in an HPC or large-scale compute environment. 3+ years of experience building and operating Ansible automation at scale (roles/collections, idempotency, inventories, secrets). Strong Linux hardening & compliance fundamentals (SELinux/AppArmor, SSH key automation, baseline config management). Demonstrated experience operating or automating clustered compute environments (HPC, large Linux farms, or similar). Hands-on experience with container tooling in Linux environments, including image lifecycle/versioning. Familiarity with incident response and runbook-driven operations; ability to automate common remediations. Strong Git workflow and documentation practices. Must hold at least one active/current technical certification from the following- -Systems engineering (e.g., INCOSE) -Information security (e.g., CISSP) -Networking (e.g., CCNA) -System Administration (e.g., RHCE, MCSE) -Virtualization (e.g., VCP) -IT systems management (e.g., ITIL) -Project management (e.g., PMP, Agile) This position requires an active/current TS/SCI w/ Polygraph. Preferred qualifications Bare-metal provisioning experience (PXE/iPXE, Kickstart/Preseed, Foreman/MAAS) and hardware OOB management. CI/testing for automation and promotion pipelines for playbooks Experience with tuned performance profiles, HPC performance troubleshooting, and GPU node health validation. Experience generation operational documentation from repos (Confluence)., TO BE CONSIDERED FOR THIS POSITION YOU MUST CURRENTLY HAVE AN ACTIVE TS/SCI WITH POLYGRAPH SECURITY CLEARANCE WITH THE FEDERAL GOVERNMENT. (U.S. CITIZENSHIP REQUIRED). Senior lev…

  • 1 month ago

Benefits & conditions

We offer a very competitive benefits package because we care about our employees. Our offerings including: Four (4) weeks paid time Eleven Paid Holidays 401k plan with 4% employer contributions and immediate vesting plus annual 3% contruibution Annual Bonuses for performance Medical, Dental, and Vision and insurance Education and Training budget provided Technical Certification and coursework budget for computer/program expenses Gym Membership Reimbursement Amazon Prime Membership Reimbursement Birthday Club Capstone Swag

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all