> Markdown version of [/jobs/ext/2707953-ai-hpc-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2707953-ai-hpc-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI & HPC Infrastructure Engineer - **Company:** Prodapt Asic Services - **Location:** Santa Clara, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Linux, Performance Tuning, Productivity Software, Virtualization Technology, AI Infrastructure, Data Logging, Cloud Platform System, High Performance Computing, Application Specific Integrated Circuits, System Availability, Large Language Models, Containerization, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Slurm, Machine Learning Operations, Engineering Base, Docker - **Published:** September 4, 2026 - **Apply:** https://www.disabledperson.com/jobs/74865763-ai-hpc-infrastructure-engineer ## About the Role * Bachelor's degree in Computer Science, Engineering, or a related technical discipline. * 3+ years of hands-on experience in Infrastructure Engineering, Platform Engineering, Cloud Operations, or HPC environments. * Strong experience with Linux-based infrastructure administration. * Experience working with public cloud platforms such as AWS, Azure, or GCP. * Hands-on experience with GPU/HPC environments and workload orchestration platforms such as Kubernetes or Slurm. * Experience with automation, infrastructure provisioning, monitoring, and performance optimization. * Solid understanding of compute, storage, networking, virtualization, and container technologies. Preferred Qualifications * Experience supporting AI/ML workloads and GPU-based infrastructure. * Knowledge of Kubernetes, Docker, Infrastructure-as-Code tools, and observability platforms. * Familiarity with AI platforms, LLM deployment, and modern engineering productivity tools. ## Description In this role, you will deploy, automate, and manage GPU-enabled infrastructure across on-premises and cloud environments, enabling scalable, reliable, and cost-effective compute resources for engineering and R&D teams. You will work at the intersection of AI infrastructure, cloud platforms, automation, and operations to support next-generation AI and engineering applications., * Build, configure, and operate GPU and HPC clusters across compute, storage, and networking environments. * Support capacity planning, performance tuning, and resource optimization for AI training, inference, and compute-intensive workloads. * Monitor infrastructure health and ensure high availability and performance. * Deploy and manage compute environments across on-premises and public cloud platforms (AWS, Azure, or GCP). * Contribute to infrastructure modernization, scalability, and resiliency initiatives. * Support cloud adoption and hybrid computing strategies. * Implement Infrastructure-as-Code (IaC) and automation frameworks for provisioning and operations. * Develop monitoring, logging, and alerting solutions to improve platform reliability. * Drive continuous improvements in operational efficiency and resource utilization. * Deploy, integrate, and support AI services, including LLM APIs, coding assistants, and AI/agent platforms. * Collaborate with engineering teams to enable AI-driven development workflows. * Support AI/ML infrastructure requirements and best practices. * Troubleshoot and resolve infrastructure, networking, and platform issues. * Create and maintain technical documentation, standards, and operational runbooks. * Partner with engineering, IT, and platform teams to deliver secure and scalable solutions. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Transforming Software Development: The Role of AI and Developer Tools](https://www.wearedevelopers.com/magazine/527-transforming-software-development-the-role-of-ai-and-developer-tools)