AI & HPC Infrastructure Engineer

Prodapt Asic Services
San Jose, CA, United States
4 days ago
Apply on www.disabledperson.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Linux Performance Tuning Productivity Software Virtualization Technology AI Infrastructure Data Logging Cloud Platform System High Performance Computing
+12 more
Application Specific Integrated Circuits System Availability Large Language Models Containerization AI Platforms Kubernetes Infrastructure Automation Frameworks Information Technology Slurm Machine Learning Operations Engineering Base Docker

Job description

In this role, you will deploy, automate, and manage GPU-enabled infrastructure across on-premises and cloud environments, enabling scalable, reliable, and cost-effective compute resources for engineering and R&D teams. You will work at the intersection of AI infrastructure, cloud platforms, automation, and operations to support next-generation AI and engineering applications., * Build, configure, and operate GPU and HPC clusters across compute, storage, and networking environments.

  • Support capacity planning, performance tuning, and resource optimization for AI training, inference, and compute-intensive workloads.
  • Monitor infrastructure health and ensure high availability and performance.
  • Deploy and manage compute environments across on-premises and public cloud platforms (AWS, Azure, or GCP).
  • Contribute to infrastructure modernization, scalability, and resiliency initiatives.
  • Support cloud adoption and hybrid computing strategies.
  • Implement Infrastructure-as-Code (IaC) and automation frameworks for provisioning and operations.
  • Develop monitoring, logging, and alerting solutions to improve platform reliability.
  • Drive continuous improvements in operational efficiency and resource utilization.
  • Deploy, integrate, and support AI services, including LLM APIs, coding assistants, and AI/agent platforms.
  • Collaborate with engineering teams to enable AI-driven development workflows.
  • Support AI/ML infrastructure requirements and best practices.
  • Troubleshoot and resolve infrastructure, networking, and platform issues.
  • Create and maintain technical documentation, standards, and operational runbooks.
  • Partner with engineering, IT, and platform teams to deliver secure and scalable solutions.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
  • 3+ years of hands-on experience in Infrastructure Engineering, Platform Engineering, Cloud Operations, or HPC environments.
  • Strong experience with Linux-based infrastructure administration.
  • Experience working with public cloud platforms such as AWS, Azure, or GCP.
  • Hands-on experience with GPU/HPC environments and workload orchestration platforms such as Kubernetes or Slurm.
  • Experience with automation, infrastructure provisioning, monitoring, and performance optimization.
  • Solid understanding of compute, storage, networking, virtualization, and container technologies.

Preferred Qualifications

  • Experience supporting AI/ML workloads and GPU-based infrastructure.
  • Knowledge of Kubernetes, Docker, Infrastructure-as-Code tools, and observability platforms.
  • Familiarity with AI platforms, LLM deployment, and modern engineering productivity tools.

About the company

Prodapt is the largest specialized player in the Connectedness industry. As an AI-first strategic technology partner, Prodapt provides consulting, business reengineering, and managed services for the largest telecom and tech enterprises building networks and digital experiences of tomorrow. A ServiceNow-invested company, Prodapt has been recognized by Gartner as a Large, Telecom-Native, Regional IT Service Provider. Prodapt’s ASIC Services is a leading provider of SoC/ASIC RTL Design, UVM based verification, Emulation, FPGA based validation, DFT, RTL2GDSII, Physical Design using ICC2 and Innovus, Mask Layout, Firmware, Silicon Bringup, and Analog mask layout. Our embedded services include device drivers, RTOS porting, and board bring-up. A “Great Place To Work® Certified “ company, Prodapt employs over 5,000 technology and domain experts in 30+ countries. Prodapt is part of the 130-year-old business conglomerate The Jhaver Group, which employs over 32,000 people across 80+ locations

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.disabledperson.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:29 min

Tech infrastructure capacity and AI product innovations

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all