> Markdown version of [/jobs/ext/2713862-senior-kubernetes-engineer-in-dallas](https://www.wearedevelopers.com/jobs/ext/2713862-senior-kubernetes-engineer-in-dallas). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Kubernetes Engineer in Dallas - **Company:** Energy Jobline - **Location:** Dallas, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Monitoring of Systems, Python (Programming Language), Machine Learning, Performance Tuning, Role-Based Access Control, Prometheus, Scientific Computating, AI Infrastructure, Cloud Platform System, High Performance Computing, Large Language Models, Grafana, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Slurm, Terraform - **Published:** September 4, 2026 - **Apply:** https://www.energyjobline.com/job/senior-kubernetes-engineer-dallas-31440644 ## About the Role * Strong experience operating Kubernetes in large-scale, production environments * Hands-on experience with NVIDIA GPU ecosystem, including GPU Operator, device plugins, MIG, and DCGM * Proficiency in Go or Python for building Kubernetes operators and automation tooling * Deep understanding of Kubernetes internals, including CRDs, controllers, RBAC, and scheduling * Experience supporting GPU-intensive workloads such as AI/ML training, LLMs, or scientific computing * Experience with GitOps, CI/CD pipelines, and infrastructure-as-code practices * Familiarity with container networking, including CNI plugins such as NVIDIA CNI or Multus * Experience with monitoring and observability tools for cluster and GPU performance This is a high-impact opportunity to work at the forefront of AI infrastructure, helping build and scale the platforms that power next- compute. ## Description We are seeking a Senior Kubernetes Engineer to help design and scale a next- GPU-accelerated compute platform supporting AI, machine learning, and high-performance computing workloads. This role sits at the core of a rapidly expanding infrastructure environment, focused on building high-throughput, highly efficient container platforms across on-prem and hybrid environments. You will play a key role in architecting and operating large-scale Kubernetes clusters optimized for GPU workloads, working closely with platform, HPC, and ML engineering teams to deliver reliable, multi-tenant compute at scale. This is a hands-on engineering role with strong ownership across performance, automation, and platform evolution., Kubernetes Platform Engineering * Design, deploy, and operate large-scale Kubernetes clusters optimized for GPU-intensive workloads * Architect container platforms supporting AI/ML, LLM training, and HPC use cases * Extend Kubernetes through custom operators, controllers, and CRDs to support infrastructure automation GPU & Workload Optimization * Integrate and optimize NVIDIA ecosystem components, including GPU Operator, DCGM, and device plugins * Implement GPU scheduling strategies, including MIG, sharing, and workload placement optimization * Enhance cluster efficiency using scheduler extensions such as kube-scheduler plugins, Slurm, or Volcano Platform Performance & Reliability * Drive performance tuning across compute, networking, and storage layers for high-throughput workloads * Partner with HPC and ML teams to ensure scalability, reliability, and workload efficiency * Participate in production readiness, incident response, and continuous improvement initiatives Observability & Automation * Implement monitoring and telemetry solutions using Prometheus, Grafana, DCGM Exporter, and OpenTelemetry * Build and maintain CI/CD pipelines for infrastructure using GitOps tools such as ArgoCD and FluxCD * Contribute to infrastructure-as-code using Terraform, Helm, and Kustomize Security & Multi-Tenancy * Design and enforce secure multi-tenant environments with namespace isolation, RBAC, and policy controls * Implement governance frameworks using tools such as OPA or Gatekeeper * Ensure compliance with platform security and operational standards ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)