TELECOMMUTE Lead HPC Kubernetes Engineer

EPAM Systems, Inc.
United States
8 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Continuous Integration Software Debugging Job Scheduling Python (Programming Language) Azure Machine Learning Graphics Processing Unit (GPU) Google Cloud AWS ECS
+3 more
Kubernetes Terraform Oracle Cloud Infrastructure

Job description

We are seeking a Lead HPC Kubernetes Engineer to help our customer develop and manage several HPC clusters across AWS, CoreWeave, Google Cloud Platform, and other providers, spanning several thousand GPUs today and scaling to 10x in 2026 and beyond. This role is Kubernetes-heavy, operating multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours, at a scale where novel failure modes are routine. Responsibilities Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers Take ownership of cluster lifecycle, node pool management, networking policy, and stability maintenance during rapid growth Provision HPC infrastructure through CI/CD systems across AWS, CoreWeave, Google Cloud Platform, and OCI, with additional providers to be added in the near future Manage job scheduling to allocate GPU compute across training and inference workloads Define and maintain SLIs/SLOs Build monitoring and

Requirements

alerting systems Participate in severity escalation response and author post-incident reviews Coordinate daily with Networking, Storage, Security, and AI/ML platform teams Requirements 5+ years of experience in infrastructure engineering, cloud platforms, or HPC Expertise in Kubernetes, with hands-on experience operating clusters at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets Proficiency in Terraform for writing and reviewing infrastructure-as-code daily Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre) Skills in Python for tooling and automation English proficiency at B2 level or higher Nice to have Familiarity with Google Kubernetes Engine Familiarity with Amazon Elastic Kubernetes Service Knowledge of Google Cloud Platform

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:52 min

Deploying Celery workers on AWS ECS Fargate containers

Jan Giacomelli · LIVE

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann Chris Heilmann +3 · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:59 min

Technical architecture and automated cloud infrastructure components

Jordi Abad · World Congress 2022

1:27 min

Scaling engineering talent through distributed remote hubs

David Singleton · World Congress 2022

Videos

See all

Related articles

See all