> Markdown version of [/jobs/ext/1578762-kubernetes-engineer](https://www.wearedevelopers.com/jobs/ext/1578762-kubernetes-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Kubernetes Engineer - **Company:** Interon IT Solutions LLC - **Location:** McLean, VA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Bash Shell, Big Data, Cloud Engineering, Computer Programming, Information Engineering, Distributed Computing Environment, Distributed Systems, Identity and Access Management, Python (Programming Language), Node.Js, Azure Machine Learning, Data Logging, Data Processing, Cloud Platform System, Autoscaling, Apache Spark, Amazon Virtual Private Cloud (VPC), Cloudformation, Kubernetes, Cloudwatch, Golang - **Published:** July 16, 2026 - **Apply:** https://www.careerjet.com/jobad/us525160db66179837d44371df9b686b70 ## About the Role * 8+ years of infrastructure, cloud engineering, or platform engineering experience. * Deep expertise administering Kubernetes in large-scale production environments. * Strong experience designing and operating Amazon EKS clusters. * Experience with Kubernetes autoscaling technologies including Karpenter or Cluster Autoscaler. * Strong understanding of Kubernetes scheduling, networking, storage, security, and cluster lifecycle management. * Experience supporting Apache Spark or other distributed compute frameworks in Kubernetes environments. * Hands-on experience with AWS services including EC2, EBS, EFS, IAM, VPC, CloudWatch, and Auto Scaling. * Experience operating highly available, production-critical systems with a focus on performance, resiliency, and automation. * Strong scripting or programming experience using Python, Go, or Bash. * Experience implementing Infrastructure as Code using Terraform, CloudFormation, or similar technologies. * Strong troubleshooting skills across distributed systems and cloud-native infrastructure. Preferred Qualifications * Kubernetes certifications (CKA, CKAD, or CKS). * AWS Certified Solutions Architect or AWS Certified Kubernetes-related certifications. * Experience operating air-gapped or highly secure cloud environments. * Contributions to Kubernetes, Karpenter, Spark, or other cloud-native open-source projects. * Experience implementing FinOps and cloud cost optimization strategies. * Background supporting large-scale data engineering, analytics, or AI/ML platforms. * Familiarity with GitOps tools such as ArgoCD or Flux., Lead Software Engineer (Python, Kubernetes) Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast-paced, collabora… + 1 month ago ## Description We are seeking a Senior Kubernetes Engineer to design, build, and optimize highly scalable Kubernetes infrastructure supporting large-scale, data-intensive workloads in AWS. This is a hands-on engineering role focused on Amazon EKS, Kubernetes platform operations, and distributed computing environments where reliability, automation, and performance are critical. The ideal candidate has deep expertise in Kubernetes internals, cluster operations, and cloud-native infrastructure, with experience supporting large-scale Apache Spark or similar distributed processing platforms. You'll work alongside platform and data engineering teams to build resilient, secure, and cost-efficient infrastructure capable of supporting thousands of concurrent workloads. What You'll Do * Design, deploy, and maintain highly available Amazon EKS clusters supporting large-scale data processing workloads. * Build and operate secure Kubernetes environments within private AWS VPCs, including air-gapped deployments, private container registries, and internal package repositories. * Troubleshoot complex Kubernetes, Karpenter, and distributed application issues including scheduling, autoscaling, networking, and cluster performance. * Optimize node provisioning using Karpenter, balancing workload performance, resiliency, and cloud cost optimization. * Design and implement strategies for Spot and On-Demand capacity management, including graceful interruption handling and workload recovery. * Configure Kubernetes resource management using ResourceQuotas, LimitRanges, PriorityClasses, taints, tolerations, and affinity rules to maximize cluster efficiency. * Deploy and optimize persistent storage solutions using Amazon EBS CSI and Amazon EFS CSI drivers for high-performance data processing workloads. * Build observability solutions with centralized logging, monitoring, alerting, and performance dashboards to proactively identify issues before production impact. * Design resilient platform architectures utilizing checkpointing, retry mechanisms, fault isolation, and automated recovery strategies. * Partner with platform, infrastructure, and data engineering teams to improve scalability, automation, security, and operational excellence., Lead Software Engineer (AWS, Golang, NodeJS) At Capital One, we are changing banking for good by creating responsible and reliable AI-powered systems. Our investments in technolo… + 15 hours ago + ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)