> Markdown version of [/jobs/ext/2266429-ai-infrastructure-engineer-l3](https://www.wearedevelopers.com/jobs/ext/2266429-ai-infrastructure-engineer-l3). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure Engineer L3 - **Company:** HCL America Inc. - **Location:** Santa Clara, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Ubuntu (Operating System), Cloud Computing, Cloud Engineering, Computer Clusters, Nvidia CUDA, Software Debugging, Linux, Distributed Computing Environment, Distributed Systems, Firmware, InfiniBand, Linux System Administration, Performance Tuning, Remote Direct Memory Access, Red Hat Enterprise Linux, Reliability Engineering, Prometheus, AI Infrastructure, Ceph (Software), High Performance Computing, System Availability, Grafana, Multi-Agent Systems, Containerization, Kubernetes, Information Technology, Slurm, Machine Learning Operations, TensorRT, Hardware Infrastructure, Terraform - **Published:** August 27, 2026 - **Apply:** https://www.dice.com/job-detail/0dfd9a5e-6a11-4f3d-8ab6-a5c1132174ab ## About the Role * Strong experience with NVIDIA GPU platforms and GPU cluster administration. * Expertise in Kubernetes, containerization, and cloud-native technologies. * Hands-on experience with CUDA, TensorRT, NCCL, DeepSpeed, Horovod, and distributed training. * Strong Linux administration and performance tuning skills. * Experience with Terraform, Helm, ArgoCD, and automation frameworks. * Knowledge of AI infrastructure, MLOps, and large-scale distributed systems. * Excellent troubleshooting, debugging, and production support experience. Preferred Certifications * NVIDIA Certified Associate AI Infrastructure * NVIDIA Base Command Manager Certification * AWS Solutions Architect Associate * Certified Kubernetes Administrator (CKA) * Certified Kubernetes Application Developer (CKAD), * Bachelor's Degree in Computer Science, Engineering, or a related field. * 8-12 years of Infrastructure or Platform Engineering experience. * 4-6 years supporting AI/ML environments and GPU-based platforms. * Experience operating production-scale AI infrastructure. ## Description We are seeking an experienced AI Infrastructure Engineer (L3) to design, deploy, optimize, and support high-performance AI and Machine Learning infrastructure. The ideal candidate will have deep expertise in GPU platforms, Kubernetes, HPC environments, distributed systems, and cloud-native AI technologies. This role involves managing large-scale GPU clusters, supporting AI training and inference workloads, troubleshooting complex infrastructure issues, and driving platform reliability., * Deploy and manage NVIDIA GPU infrastructure (A100, H100, L40) and AI accelerator platforms. * Administer Kubernetes GPU clusters using NVIDIA GPU Operator and related technologies. * Install and maintain CUDA, cuDNN, TensorRT, firmware, and driver stacks. * Manage high-performance storage solutions such as Ceph, Lustre, BeeGFS, and NFS. * Support InfiniBand, RDMA, RoCE, NVLink, and other high-speed networking technologies. * Optimize Linux environments (RHEL, Ubuntu, Rocky Linux) for AI and HPC workloads. * Support AI orchestration platforms including Kubeflow, MLflow, Ray, and Slurm. * Implement Infrastructure as Code using Terraform, Helm, and GitOps tools. * Monitor platform performance with Prometheus, Grafana, NVIDIA DCGM, and OpenTelemetry. * Lead root cause analysis (RCA) and resolve critical GPU, networking, storage, and platform issues. * Collaborate with cloud, data science, MLOps, SRE, and engineering teams to deliver scalable AI platforms. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)