Kubernetes Engineer With HPC Engineer

BURGEON IT SERVICES LLC
United States
5 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Continuous Integration Software Debugging Job Scheduling Python (Programming Language) Azure Machine Learning Google Cloud Multi-Cloud Kubernetes
+2 more
Terraform Oracle Cloud Infrastructure

Job description

  • This role is Kubernetes-heavy. You’ll operate multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours. The clusters are large enough that novel failure modes are routine., * Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers. You’re responsible for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth.
  • Provision HPC infrastructure through CI/CD system across AWS, CoreWeave, Google Cloud Platform, and OCI, with additional providers to be expanded in the near future.
  • Manage job scheduling to allocate GPU compute across training and inference workloads.
  • Define and maintain SLIs/SLOs. Build monitoring and alerting. Participate in severity escalation response and author post-incident reviews.
  • Coordinate daily with Networking, Storage, Security, and AI/ML platform teams.

Requirements

  • 4+ years in infrastructure engineering, cloud platforms, or HPC.
  • Kubernetes is the core requirement. You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit.
  • Terraform proficiency. You’ll write and review infrastructure-as-code daily.
  • Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre).
  • Python for tooling and automation.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:43 min

Bursting GPU capacity over hybrid Kubernetes networks

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann Chris Heilmann +3 · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:15 min

Deploying local container pods to Kubernetes clusters

Stevan Le Meur Stevan Le Meur · World Congress 2024

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:44 min

Automating storage savings with S3 intelligent tiering

Sébastien Stormacq · World Congress 2021

Videos

See all

Related articles

See all