Kubernetes Engineer

Hptech Inc.
Dallas, TX, United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Big Data Information Engineering Software Debugging File Systems Distributed Systems Fault Tolerance Open Source Technology Data Logging Data Processing System Availability Apache Spark
+3 more
Amazon Virtual Private Cloud (VPC) Kubernetes Block Storage

Job description

  • Apache Spark optimization to support thousands of concurrent jobs processing large volumes of Bigdata. (candidates are working in smaller environments)
  • Active contributions to open-source Kubernetes, Karpenter, or Spark projects
  • Configure persistent volume claims and provision container-native storage using Amazon EBS CSI driver for block storage and Amazon EFS CSI driver for shared file access, We are seeking an expert Kubernetes engineer to configure, deploy, and operate mission-critical Amazon EKS infrastructure supporting petabyte-scale data processing workloads. This role requires deep technical expertise in Kubernetes internals, large-scale cluster management, and Apache Spark optimization to support thousands of concurrent jobs processing large volumes of Bigdata.

Core Responsibilities

  • Design, deploy, and maintain production-grade Amazon EKS clusters architected for petabyte-scale data processing with high availability and fault tolerance
  • Operate in air-gapped private VPC environments without internet access, managing secure package repositories and container registries
  • Debug and resolve complex distributed systems issues across EKS, Karpenter, and Spark, including scheduling bottlenecks, node scaling delays, and cascading failures under heavy load
  • Implement Karpenter consolidation and disruption policies balancing cost optimization with job resiliency
  • Manage spot and on-demand instance strategies with robust node interruption handling
  • Define and enforce ResourceQuotas, LimitRanges, and PriorityClasses to ensure fair resource distribution and prevent resource starvation
  • Configure persistent volume claims and provision container-native storage using Amazon EBS CSI driver for block storage and Amazon EFS CSI driver for shared file access
  • Optimize storage configurations for resiliency, high-throughput data processing workloads
  • Implement logging, alerting, and anomaly detection to identify provisioning failures, executor loss, and throughput degradation before broader system impact
  • Design fault-tolerant architectures with retry strategies, checkpointing, and graceful degradation patterns minimizing re-computation on failure
  • Develop and manage configs in a private VPC without internet access

Requirements

  • Kubernetes certification (CKA, CKAD, or CKS) & AWS certifications
  • Active contributions to open-source Kubernetes, Karpenter, or Spark projects
  • Familiarity with FinOps practices and cost optimization at scale
  • Background in data engineering or analytics platforms

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

3:40 min

Accessing local files via the File System Access API

Rowdy Rabouw Rowdy Rabouw · WWC 2025

1:54 min

Speaker background and open source Kubernetes edge computing projects

Gaurav Gahlot Gaurav Gahlot · WWC Europe 2026

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all