Senior DevOps Engineer (HPC / EDA / SLURM)
Spectraforce
Rancho Cordova, CA, United States
2 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source
Tech stack
Active Directory
Confluence
JIRA
User Authentication
Bash Shell
Code Review
System Configuration
Data Centers
Software Debugging
Linux
DevOps
File Systems
+31 more
Perl (Programming Language)
VMware ESX Servers
Github
Identity and Access Management
Python (Programming Language)
Lightweight Directory Access Protocols (LDAP)
NetApp Applications
Red Hat Enterprise Linux
Migration Manager
Ansible
Runbook
Scientific Computating
SUSE Linux Enterprise Servers
Software Configuration Management
Virtual Machines
VMware VSphere
YAML
Zabbix
Data Logging
Software Repository
Scripting
High Performance Computing
Okta
Git
Bare Metal
Data Management
Slurm
Splunk
Software Version Control
Servicenow
Artifactory
Job description
HPC / EDA Platform Operations
- Support and administer SLURM-based HPC compute environments, including partition configuration and migration planning
- Plan and execute HPC/EDA compute and storage infrastructure migrations across datacenters
- Develop migration strategies and evaluate implementation options, risks, dependencies, and operational tradeoffs
- Author formal Method of Procedure (MOP) documents and runbooks for infrastructure changes and service cutovers
- Coordinate cross-functionally with EDA/SPE teams, storage teams, and IDAM to deliver coordinated platform changes
- Verify storage volumes, application access, and service continuity following migrations or infrastructure changes
- Define HPC storage service tiers and gather performance and capacity requirements for EDA workloads
Linux Systems Engineering & OS Deployment
- Administer SUSE Linux Enterprise Server (SLES) 12 and SLES 15 systems in a production HPC environment
- Build and maintain custom Linux OS images and installation media using Kiwi NG and related tooling
- Enable and maintain bare-metal provisioning workflows via RackN / Digital Rebar Provision for SLES and ESXi deployments add PXE
- Provision and configure VMware vSphere virtual machines for HPC service workloads
- Troubleshoot production issues in Linux HPC operational scripts, services, and system daemons, * Develop, maintain, and extend Ansible playbooks and roles for Linux system setup, authentication, and platform configuration
- Ensure multi-version Ansible playbook compatibility across SLES 12 and SLES 15
- Manage Git repositories and Artifactory artifact storage; migrate large binaries and configuration artifacts out of source control
- Contribute GitHub pull requests, conduct code reviews, and manage inner-source infrastructure repositories
- Drive production environment changes through change management workflows using ServiceNow
- Identity & Access Management
- Integrate and configure enterprise identity systems including Okta, Active Directory, LDAP, NIS, VAS, and SSSD for Linux/HPC environments
- Audit and reconcile Linux user and group identity data (UID/GID) across multiple directory and authentication domains
- Validate authentication methods and access behavior across HPC compute and storage environments
- Extend SSSD-based corporate authentication to new compute environments and author corresponding Ansible automation
Monitoring, Logging & Operational Readiness
- Assess and implement log management strategies, including evaluation of Splunk integration for HPC system logs
- Investigate and remediate operational issues in production Linux services (VNC, NIS, AutoFS, Zabbix, etc.)
- Produce technical documentation, architecture diagrams, implementation guides, and end-user instructions in Confluence
Requirements
- HPC / EDA Platforms: SLURM, HPC compute/storage administration, EDA infrastructure, datacenter migrations
- Linux / OS: SLES 12, SLES 15, ESXi 8.0, Kiwi NG ISO creation, Linux system services, acct/pacct, dracut
- Provisioning / Automation: Ansible (playbooks, roles, multi-version), RackN / Digital Rebar Provision, VMware vSphere
- Identity / Auth: SSSD, Okta, Active Directory, LDAP, NIS, VAS, UID/GID auditing, cross-domain identity management
- Storage / Filesystems: NetApp SVM, NFS, AutoFS, RootSquash, storage tier design, IOPS/capacity planning
- DevOps / Source Control: Git, GitHub, Artifactory, inner-source repository management
- Monitoring / Logging: Splunk integration, Zabbix, operational script hardening, log management
- Scripting / Languages; Ansible (YAML), Python, Perl (debugging), Bash
- ITSM / Documentation: ServiceNow (change requests), MOP authoring, Confluence, Jira, technical diagramming
Experience Requirements
- 5+ years of experience in a DevOps, Platform Engineering, or Linux Systems Engineering role
- Hands-on HPC cluster administration experience, including SLURM or equivalent workload managers
- Demonstrated experience supporting EDA or scientific computing environments
- Strong Ansible automation skills with production-grade playbook and role development
- Experience with bare-metal provisioning tools (RackN, Cobbler, or equivalent)
- Proven ability to plan and execute datacenter or infrastructure migrations with minimal disruption
- Familiarity with enterprise Linux identity and authentication stacks (SSSD, LDAP, AD, NIS, Okta)
- Experience with NetApp or comparable enterprise storage platforms in HPC contexts
- Ability to author formal technical documentation (MOPs, runbooks, architecture diagrams)
- Strong written and verbal communication skills; capable of coordinating across multiple teams
Preferred Qualifications
- Experience with SUSE Linux Enterprise Server (SLES) 12 and/or 15 in an enterprise environment
- Familiarity with RackN / Digital Rebar Provision for bare-metal OS deployment
- Hands-on experience with Kiwi NG or similar tools for custom OS image creation
- Knowledge of VMware vSphere for HPC support VM provisioning
- Experience migrating configuration artifacts and binaries to Artifactory
- Background in semiconductor, storage, or high-tech manufacturing IT environments
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on leoforce.usGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
EM
Eli McGarvie
DevOps Engineer Salary [2023]
over 3 years ago
LM
Luis Minvielle
Top 6 Hackathons for Developers in 2023
about 3 years ago
JF
Jonas Fritzsch
Résumé-Driven Development: How IT trends affect the job market for software developers
almost 5 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
about 2 years ago