> Markdown version of [/jobs/ext/1928690-kubernetes-engineer](https://www.wearedevelopers.com/jobs/ext/1928690-kubernetes-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Kubernetes Engineer - **Company:** Hptech Inc. - **Location:** Dallas, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Big Data, Information Engineering, Software Debugging, File Systems, Distributed Systems, Fault Tolerance, Open Source Technology, Data Logging, Data Processing, System Availability, Apache Spark, Amazon Virtual Private Cloud (VPC), Kubernetes, Block Storage - **Published:** August 5, 2026 - **Apply:** https://www.dice.com/job-detail/1c7126f1-cbf1-488a-b5df-dbb4182db67b ## About the Role * Kubernetes certification (CKA, CKAD, or CKS) & AWS certifications * Active contributions to open-source Kubernetes, Karpenter, or Spark projects * Familiarity with FinOps practices and cost optimization at scale * Background in data engineering or analytics platforms ## Description * Apache Spark optimization to support thousands of concurrent jobs processing large volumes of Bigdata. (candidates are working in smaller environments) * Active contributions to open-source Kubernetes, Karpenter, or Spark projects * Configure persistent volume claims and provision container-native storage using Amazon EBS CSI driver for block storage and Amazon EFS CSI driver for shared file access, We are seeking an expert Kubernetes engineer to configure, deploy, and operate mission-critical Amazon EKS infrastructure supporting petabyte-scale data processing workloads. This role requires deep technical expertise in Kubernetes internals, large-scale cluster management, and Apache Spark optimization to support thousands of concurrent jobs processing large volumes of Bigdata. Core Responsibilities * Design, deploy, and maintain production-grade Amazon EKS clusters architected for petabyte-scale data processing with high availability and fault tolerance * Operate in air-gapped private VPC environments without internet access, managing secure package repositories and container registries * Debug and resolve complex distributed systems issues across EKS, Karpenter, and Spark, including scheduling bottlenecks, node scaling delays, and cascading failures under heavy load * Implement Karpenter consolidation and disruption policies balancing cost optimization with job resiliency * Manage spot and on-demand instance strategies with robust node interruption handling * Define and enforce ResourceQuotas, LimitRanges, and PriorityClasses to ensure fair resource distribution and prevent resource starvation * Configure persistent volume claims and provision container-native storage using Amazon EBS CSI driver for block storage and Amazon EFS CSI driver for shared file access * Optimize storage configurations for resiliency, high-throughput data processing workloads * Implement logging, alerting, and anomaly detection to identify provisioning failures, executor loss, and throughput degradation before broader system impact * Design fault-tolerant architectures with retry strategies, checkpointing, and graceful degradation patterns minimizing re-computation on failure * Develop and manage configs in a private VPC without internet access ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Possibilities with Web Capabilities](https://www.wearedevelopers.com/videos/1575-possibilities-with-web-capabilities) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [Developing locally with Kubernetes - a Guide and Best Practices](https://www.wearedevelopers.com/videos/851-developing-locally-with-kubernetes-a-guide-and-best-practices) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)