> Markdown version of [/jobs/ext/3093609-telecommute-lead-hpc-kubernetes-engineer](https://www.wearedevelopers.com/jobs/ext/3093609-telecommute-lead-hpc-kubernetes-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # TELECOMMUTE Lead HPC Kubernetes Engineer - **Company:** EPAM Systems, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Continuous Integration, Software Debugging, Job Scheduling, Python (Programming Language), Azure Machine Learning, Graphics Processing Unit (GPU), Google Cloud, AWS ECS, Kubernetes, Terraform, Oracle Cloud Infrastructure - **Published:** September 26, 2026 - **Apply:** https://www.dice.com/job-detail/45e1bd91-b5d9-4ea4-9b18-835cf63413a1 ## About the Role alerting systems Participate in severity escalation response and author post-incident reviews Coordinate daily with Networking, Storage, Security, and AI/ML platform teams Requirements 5+ years of experience in infrastructure engineering, cloud platforms, or HPC Expertise in Kubernetes, with hands-on experience operating clusters at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets Proficiency in Terraform for writing and reviewing infrastructure-as-code daily Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre) Skills in Python for tooling and automation English proficiency at B2 level or higher Nice to have Familiarity with Google Kubernetes Engine Familiarity with Amazon Elastic Kubernetes Service Knowledge of Google Cloud Platform ## Description We are seeking a Lead HPC Kubernetes Engineer to help our customer develop and manage several HPC clusters across AWS, CoreWeave, Google Cloud Platform, and other providers, spanning several thousand GPUs today and scaling to 10x in 2026 and beyond. This role is Kubernetes-heavy, operating multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours, at a scale where novel failure modes are routine. Responsibilities Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers Take ownership of cluster lifecycle, node pool management, networking policy, and stability maintenance during rapid growth Provision HPC infrastructure through CI/CD systems across AWS, CoreWeave, Google Cloud Platform, and OCI, with additional providers to be added in the near future Manage job scheduling to allocate GPU compute across training and inference workloads Define and maintain SLIs/SLOs Build monitoring and ## Related Videos - [Celery on AWS ECS - the art of background tasks & continuous deployment](https://www.wearedevelopers.com/videos/561-celery-on-aws-ecs-the-art-of-background-tasks-continuous-deployment) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)