Kubernetes Engineer
Hptech Inc.
Dallas, TX, United States
6 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Big Data
Information Engineering
Software Debugging
File Systems
Distributed Systems
Fault Tolerance
Open Source Technology
Data Logging
Data Processing
System Availability
Apache Spark
+3 more
Amazon Virtual Private Cloud (VPC)
Kubernetes
Block Storage
Job description
- Apache Spark optimization to support thousands of concurrent jobs processing large volumes of Bigdata. (candidates are working in smaller environments)
- Active contributions to open-source Kubernetes, Karpenter, or Spark projects
- Configure persistent volume claims and provision container-native storage using Amazon EBS CSI driver for block storage and Amazon EFS CSI driver for shared file access, We are seeking an expert Kubernetes engineer to configure, deploy, and operate mission-critical Amazon EKS infrastructure supporting petabyte-scale data processing workloads. This role requires deep technical expertise in Kubernetes internals, large-scale cluster management, and Apache Spark optimization to support thousands of concurrent jobs processing large volumes of Bigdata.
Core Responsibilities
- Design, deploy, and maintain production-grade Amazon EKS clusters architected for petabyte-scale data processing with high availability and fault tolerance
- Operate in air-gapped private VPC environments without internet access, managing secure package repositories and container registries
- Debug and resolve complex distributed systems issues across EKS, Karpenter, and Spark, including scheduling bottlenecks, node scaling delays, and cascading failures under heavy load
- Implement Karpenter consolidation and disruption policies balancing cost optimization with job resiliency
- Manage spot and on-demand instance strategies with robust node interruption handling
- Define and enforce ResourceQuotas, LimitRanges, and PriorityClasses to ensure fair resource distribution and prevent resource starvation
- Configure persistent volume claims and provision container-native storage using Amazon EBS CSI driver for block storage and Amazon EFS CSI driver for shared file access
- Optimize storage configurations for resiliency, high-throughput data processing workloads
- Implement logging, alerting, and anomaly detection to identify provisioning failures, executor loss, and throughput degradation before broader system impact
- Design fault-tolerant architectures with retry strategies, checkpointing, and graceful degradation patterns minimizing re-computation on failure
- Develop and manage configs in a private VPC without internet access
Requirements
- Kubernetes certification (CKA, CKAD, or CKS) & AWS certifications
- Active contributions to open-source Kubernetes, Karpenter, or Spark projects
- Familiarity with FinOps practices and cost optimization at scale
- Background in data engineering or analytics platforms
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
LM
Luis Minvielle
Top 6 Hackathons for Developers in 2023
about 3 years ago
CS
Christina Schaireiter
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
2 months ago
DS
Dhannush Subramani
Top Big Data Technologies That You Need to Know
about 4 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago