Kubernetes Engineer With HPC Engineer
BURGEON IT SERVICES LLC
United States
5 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Amazon Web Services
Amazon Elastic Compute Cloud
Amazon S3
Continuous Integration
Software Debugging
Job Scheduling
Python (Programming Language)
Azure Machine Learning
Google Cloud
Multi-Cloud
Kubernetes
+2 more
Terraform
Oracle Cloud Infrastructure
Job description
- This role is Kubernetes-heavy. You’ll operate multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours. The clusters are large enough that novel failure modes are routine., * Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers. You’re responsible for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth.
- Provision HPC infrastructure through CI/CD system across AWS, CoreWeave, Google Cloud Platform, and OCI, with additional providers to be expanded in the near future.
- Manage job scheduling to allocate GPU compute across training and inference workloads.
- Define and maintain SLIs/SLOs. Build monitoring and alerting. Participate in severity escalation response and author post-incident reviews.
- Coordinate daily with Networking, Storage, Security, and AI/ML platform teams.
Requirements
- 4+ years in infrastructure engineering, cloud platforms, or HPC.
- Kubernetes is the core requirement. You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit.
- Terraform proficiency. You’ll write and review infrastructure-as-code daily.
- Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre).
- Python for tooling and automation.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
almost 3 years ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago
CS
Christina Schaireiter
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
4 months ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
about 2 months ago
LM
Luis Minvielle
Top 6 Hackathons for Developers in 2023
about 3 years ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago