> Markdown version of [/jobs/ext/2211517-hpc-hardware-engineer](https://www.wearedevelopers.com/jobs/ext/2211517-hpc-hardware-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC Hardware Engineer - **Company:** NorthMark Strategies - **Location:** Dallas, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Intelligent Platform Management Interface, Bash Shell, BIOS, Computer Engineering, Microprocessors, Network Interface Controllers, Firmware, IBM Hardware Management Console, Monitoring of Systems, Python (Programming Language), Linux System Administration, OpenStack, Performance Tuning, Windows PowerShell, Ansible, Software Engineering, Scripting, Graphics Processing Unit (GPU), High Performance Computing, Infrastructure as Code (IaC), AI Platforms, Infrastructure Automation Frameworks, Bare Metal, Hardware Infrastructure, Puppet - **Published:** August 24, 2026 - **Apply:** https://www.dice.com/job-detail/03f31020-1de9-4b9e-aecb-1cd6296801ba ## About the Role * Bachelor's degree in Electrical Engineering, Computer Engineering, or a related field, or equivalent hands-on experience. * 5+ years of experience managing large-scale HPC or AI compute infrastructure in a production environment. * Deep knowledge of server hardware architecture, including processors, memory, storage, networking, power systems, and thermal management. * Hands-on experience with bare-metal provisioning, firmware and BIOS lifecycle management, and hardware automation tools such as Ansible, Puppet, or Chef. * Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management and diagnostics. * Demonstrated ability to troubleshoot and resolve complex hardware issues across GPU and CPU nodes, including NVIDIA-SMI and GPU diagnostics. * Experience with hardware monitoring platforms, performance tuning, and capacity planning at scale. * Familiarity with Linux-based environments and scripting proficiency in Python, Bash, or PowerShell for infrastructure automation. * Experience with OpenStack (particularly Ironic) or equivalent cloud/bare-metal provisioning platforms is strongly preferred. * Strong cross-functional communication skills and proven ability to collaborate effectively with software, networking, and vendor teams. * Prior technical leadership experience, including mentoring engineers and driving team-wide best practices. It is impossible to list every requirement for, or responsibility of, any position. Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company's needs may change over time. Therefore, the above job description is not comprehensive or exhaustive. The Company reserves the right to adjust, add to or eliminate any aspect of the above description. The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs. Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future. ## Description NMC is seeking a Senior HPC Hardware Engineer to join the Compute Engineering team based at our Dallas, TX offices at Victory Commons. This is a hands-on role at the center of one of the most demanding and rapidly scaling HPC environments in operation, spanning a large fleet of GPU and CPU nodes built on the latest NVIDIA platforms including H200 and GB200/NVL72 architectures. You will own the full hardware lifecycle for NMC 's compute fleet - from bare-metal provisioning and firmware baseline management through production validation, troubleshooting, and capacity planning. You will be the go-to expert on server hardware architecture, driving standards and automation that ensure the fleet operates at peak performance and availability. Your work directly enables the research and delivery workloads that NMC 's clients depend on. This role requires close collaboration with Software Engineering, Networking, and Vendor teams, and involves mentoring junior engineers. The ideal candidate is a technically deep infrastructure leader who thrives in fast-paced environments, brings a strong automation mindset, and has a proven track record managing large-scale HPC or AI compute infrastructure., * Design, configure, and manage a high-performance compute fleet comprising large-scale GPU (NVIDIA V100/A100/H200/GB200) and CPU nodes across NMC 's Dallas infrastructure. * Own the full firmware and BIOS lifecycle across the HPC/AI fleet - from establishing baselines and validation through rollout, compliance, and ongoing maintenance. * Lead troubleshooting of hardware components including CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs; drive component replacement and configuration remediation. * Automate health checks, onboarding workflows, and recurring hardware issue remediation to accelerate safe deployment and reduce recovery time. * Validate and operationalize next-generation AI platforms (e.g., NVL72 / Grace Blackwell) from day one, ensuring stability, performance readiness, and production fitness. * Collaborate with vendors on firmware and hardware issues, providing clear reproduction cases, diagnostic logs, and business impact to drive timely resolution. * Perform hardware performance analysis, tuning, and capacity planning to ensure reliable scale-out of the compute environment. * Define and implement security hardening best practices for hardware infrastructure, maintaining platform integrity across the fleet. * Leverage Infrastructure as Code (IaC) methodologies and scripting to drive efficient, repeatable, and scalable infrastructure management. * Mentor junior engineers, act as a subject matter expert for infrastructure-related escalations, and champion a culture of continuous improvement across the team. ## Related Videos - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [How to Submit CFPs and Get into Public Speaking - Moran Weber](https://www.wearedevelopers.com/videos/2112-how-to-submit-cfps-and-get-into-public-speaking-moran-weber) - [What if your HR software adapted to you, not the other way around?](https://www.wearedevelopers.com/videos/100259-what-if-your-hr-software-adapted-to-you-not-the-other-way-around) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)