Senior Kubernetes Engineer - Local to DMV Areas - F2f Interview
Role details
Job location
Tech stack
Job description
We are seeking a Senior Kubernetes Engineer to design, build, and optimize highly scalable Kubernetes infrastructure supporting large-scale, data-intensive workloads in AWS. This is a hands-on engineering role focused on Amazon EKS, Kubernetes platform operations, and distributed computing environments where reliability, automation, and performance are critical. The ideal candidate has deep expertise in Kubernetes internals, cluster operations, and cloud-native infrastructure, with experience supporting large-scale Apache Spark or similar distributed processing platforms. You'll work alongside platform and data engineering teams to build resilient, secure, and cost-efficient infrastructure capable of supporting thousands of concurrent workloads. What You'll Do
- Design, deploy, and maintain highly available Amazon EKS clusters supporting large-scale data processing workloads.
- Build and operate secure Kubernetes environments within private AWS VPCs, including air-gapped deployments, private container registries, and internal package repositories.
- Troubleshoot complex Kubernetes, Karpenter, and distributed application issues including scheduling, autoscaling, networking, and cluster performance.
- Optimize node provisioning using Karpenter, balancing workload performance, resiliency, and cloud cost optimization.
- Design and implement strategies for Spot and On-Demand capacity management, including graceful interruption handling and workload recovery.
- Configure Kubernetes resource management using ResourceQuotas, LimitRanges, PriorityClasses, taints, tolerations, and affinity rules to maximize cluster efficiency.
- Deploy and optimize persistent storage solutions using Amazon EBS CSI and Amazon EFS CSI drivers for high-performance data processing workloads.
- Build observability solutions with centralized logging, monitoring, alerting, and performance dashboards to proactively identify issues before production impact.
- Design resilient platform architectures utilizing checkpointing, retry mechanisms, fault isolation, and automated recovery strategies.
- Partner with platform, infrastructure, and data engineering teams to improve scalability, automation, security, and operational excellence.
Requirements
- 8+ years of infrastructure, cloud engineering, or platform engineering experience.
- Deep expertise administering Kubernetes in large-scale production environments.
- Strong experience designing and operating Amazon EKS clusters.
- Experience with Kubernetes autoscaling technologies including Karpenter or Cluster Autoscaler.
- Strong understanding of Kubernetes scheduling, networking, storage, security, and cluster lifecycle management.
- Experience supporting Apache Spark or other distributed compute frameworks in Kubernetes environments.
- Hands-on experience with AWS services including EC2, EBS, EFS, IAM, VPC, CloudWatch, and Auto Scaling.
- Experience operating highly available, production-critical systems with a focus on performance, resiliency, and automation.
- Strong scripting or programming experience using Python, Go, or Bash.
- Experience implementing Infrastructure as Code using Terraform, CloudFormation, or similar technologies.
- Strong troubleshooting skills across distributed systems and cloud-native infrastructure.
Preferred Qualifications
- Kubernetes certifications (CKA, CKAD, or CKS).
- AWS Certified Solutions Architect or AWS Certified Kubernetes-related certifications.
- Experience operating air-gapped or highly secure cloud environments.
- Contributions to Kubernetes, Karpenter, Spark, or other cloud-native open-source projects.
- Experience implementing FinOps and cloud cost optimization strategies.
- Background supporting large-scale data engineering, analytics, or AI/ML platforms.
- Familiarity with GitOps tools such as ArgoCD or Flux.
Best Regards