> Markdown version of [/jobs/ext/2459079-hpc-solutions-engineer](https://www.wearedevelopers.com/jobs/ext/2459079-hpc-solutions-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC Solutions Engineer - **Company:** HYDRA HOST, INC. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Computer Clusters, System Configuration, Distributed Systems, General-Purpose Computing on Graphics Processing Units, InfiniBand, Job Scheduling, Python (Programming Language), Machine Learning, Ansible, Systems Architecture, Graphics Processing Unit (GPU), High Performance Computing, Computer Network Technologies, Infrastructure Automation Frameworks, Slurm, Software Coding, Terraform, Software Library - **Published:** August 6, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=bfbbe519fa3faa52 ## About the Role * 7+ years of high performance compute, distributed machine learning, GPU computing, and/or system architecture experience. * Proficient in managing NVIDIA GPU environments and a familiarity with GPU computing frameworks and libraries. * Strong experience with high-speed networking technologies, specifically InfiniBand. * Experience with HPC job schedulers, preferably SLURM. * Expertise in automating environment setup and maintenance using Ansible and Terraform. * Demonstrated ability in deploying and managing virtual storage solutions. * Strong coding skills in Python and familiarity with machine learning libraries and frameworks. * Excellent problem-solving, communication, and teamwork skills. Personal Skills: * A strong communicator who is empathetic, personable and agreeable * Conscientious, meticulous, organized, and self-motivated, with the ability to define goals and prioritize workloads. * An effective team player with the ability to take ownership with a results-oriented growth mindset ## Description As a High Performance Compute Solutions Engineer on our team you will report directly to the Co-Founder & CTO and work collaboratively with other team members. You will be responsible for bespoke scoping and execution of high end compute resource orchestration for premium clients with the mission of configuring and maintaining large scale GPU clusters to best serve them. This includes installing, configuring, and monitoring dependencies required by customers and working with a wide variety of teams and workloads to ensure they can effectively leverage their large scale clusters., * Work with customers in technical discovery to help define requirements and deliverables for their use cases and help them to effectively utilize distributed GPU computing resources. * Identify and recommend the best tools for each customer, while building boilerplate and reference implementations that can be reused for subsequent customers. * Manage GPU clusters and coordinate with IT to ensure efficient operation of NVIDIA GPUs, utilizing technologies such as InfiniBand for high-speed networking. * Oversee the deployment and maintenance of machine learning environments using virtual storage solutions and distributed computing (HPC) tools such as SLURM. * Automate infrastructure provisioning and management using tools such as Ansible and Terraform. * Collaborate with data scientists and engineers to ensure seamless integration of ML models into production environments. * Conduct performance tuning and optimization of systems to maximize throughput and reduce latency. * Stay current with the latest industry trends in machine learning technologies and HPC to ensure the use of best practices in infrastructure setup and model development. * Document and maintain operational procedures and system configurations. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) - [Embracing the Hybrid Cloud: Unlocking Success with Ansible](https://www.wearedevelopers.com/videos/932-embracing-the-hybrid-cloud-unlocking-success-with-ansible) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)