AI Infrastructure Engineer L3

HCL America Inc.
Santa Clara, CA, United States
8 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Ubuntu (Operating System) Cloud Computing Cloud Engineering Computer Clusters Nvidia CUDA Software Debugging Linux Distributed Computing Environment Distributed Systems Firmware
+21 more
InfiniBand Linux System Administration Performance Tuning Remote Direct Memory Access Red Hat Enterprise Linux Reliability Engineering Prometheus AI Infrastructure Ceph (Software) High Performance Computing System Availability Grafana Multi-Agent Systems Containerization Kubernetes Information Technology Slurm Machine Learning Operations TensorRT Hardware Infrastructure Terraform

Job description

We are seeking an experienced AI Infrastructure Engineer (L3) to design, deploy, optimize, and support high-performance AI and Machine Learning infrastructure. The ideal candidate will have deep expertise in GPU platforms, Kubernetes, HPC environments, distributed systems, and cloud-native AI technologies. This role involves managing large-scale GPU clusters, supporting AI training and inference workloads, troubleshooting complex infrastructure issues, and driving platform reliability., * Deploy and manage NVIDIA GPU infrastructure (A100, H100, L40) and AI accelerator platforms.

  • Administer Kubernetes GPU clusters using NVIDIA GPU Operator and related technologies.
  • Install and maintain CUDA, cuDNN, TensorRT, firmware, and driver stacks.
  • Manage high-performance storage solutions such as Ceph, Lustre, BeeGFS, and NFS.
  • Support InfiniBand, RDMA, RoCE, NVLink, and other high-speed networking technologies.
  • Optimize Linux environments (RHEL, Ubuntu, Rocky Linux) for AI and HPC workloads.
  • Support AI orchestration platforms including Kubeflow, MLflow, Ray, and Slurm.
  • Implement Infrastructure as Code using Terraform, Helm, and GitOps tools.
  • Monitor platform performance with Prometheus, Grafana, NVIDIA DCGM, and OpenTelemetry.
  • Lead root cause analysis (RCA) and resolve critical GPU, networking, storage, and platform issues.
  • Collaborate with cloud, data science, MLOps, SRE, and engineering teams to deliver scalable AI platforms.

Requirements

  • Strong experience with NVIDIA GPU platforms and GPU cluster administration.
  • Expertise in Kubernetes, containerization, and cloud-native technologies.
  • Hands-on experience with CUDA, TensorRT, NCCL, DeepSpeed, Horovod, and distributed training.
  • Strong Linux administration and performance tuning skills.
  • Experience with Terraform, Helm, ArgoCD, and automation frameworks.
  • Knowledge of AI infrastructure, MLOps, and large-scale distributed systems.
  • Excellent troubleshooting, debugging, and production support experience.

Preferred Certifications

  • NVIDIA Certified Associate AI Infrastructure
  • NVIDIA Base Command Manager Certification
  • AWS Solutions Architect Associate
  • Certified Kubernetes Administrator (CKA)
  • Certified Kubernetes Application Developer (CKAD), * Bachelor’s Degree in Computer Science, Engineering, or a related field.
  • 8-12 years of Infrastructure or Platform Engineering experience.
  • 4-6 years supporting AI/ML environments and GPU-based platforms.
  • Experience operating production-scale AI infrastructure.

Benefits & conditions

A candidate s pay within the range will depend on their work location, skills, experience, education, and other factors permitted by law. This role may also be eligible for performance-based bonuses subject to company policies. In addition, this role is eligible for the following benefits subject to company policies: medical, dental, vision, pharmacy, life, accidental death & dismemberment, and disability insurance; employee assistance program; 401(k) retirement plan; 10 days of paid time off per year (some positions are eligible for need-based leave with no designated number of leave days per year); and 10 paid holidays per year.

About the company

HCLTech is a global technology company with over 220,000 professionals across 60 countries, delivering industry-leading capabilities in Digital, Engineering, Cloud, and AI. We help enterprises accelerate innovation through cutting-edge technologies and world-class talent.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:58 min

Verifying hardware access and exploring AI inference scaling

Piotr Zaniewski Piotr Zaniewski · World Congress 2026 Europe

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all