> Markdown version of [/jobs/ext/1318418-kubernetes-engineer](https://www.wearedevelopers.com/jobs/ext/1318418-kubernetes-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Kubernetes Engineer - **Company:** NorthMark Strategies - **Location:** Dallas, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, DevOps, Python (Programming Language), Performance Tuning, Role-Based Access Control, Prometheus, Scientific Computating, Systems Integration, Large Language Models, Grafana, Containerization, Kubernetes, Slurm, Terraform - **Published:** July 17, 2026 - **Apply:** https://www.dice.com/job-detail/f9fe5329-fed2-4573-9fe6-8642ad8c900d ## About the Role * Extensive experience with Kubernetes in production-grade environments and working with NVIDIA and Kubernetes, including GPU Operator, device plugin, NVML, MIG and DCGM * Proficiency in Go or Python for operator development and Kubernetes controller logic * Deep understanding of Kubernetes internals, including CRDs, RBAC, custom controllers and scheduler extensions * Experience with GPU-intensive workloads, for example for LLMs, training pipelines and scientific computing * Hands-on experience with Helm, Kustomize and GitOps workflows * Familiarity with CNI plugins, especially NVIDIA CNI and Multus * Experience with monitoring GPU metrics and cluster health using Prometheus and DCGM Exporter It is impossible to list every requirement for, or responsibility of, any position. Similarly, we cannot identify all the skills a position may require since job responsibilities and the Company's needs may change over time. Therefore, the above job description is not comprehensive or exhaustive. The Company reserves the right to adjust, add to or eliminate any aspect of the above description. The Company also retains the right to require all employees to undertake additional or different job responsibilities when necessary to meet business needs. Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future. ## Description In this role, you will design, implement, and optimise GPU-accelerated container platforms at scale, enabling high-performance workloads (AI/ML, HPC, LLM training) across hybrid or on-prem environments. You will have deep expertise with both NVIDIA and Kubernetes ecosystems, including GPU scheduling, device plugins and custom operators. Responsibilities * Architecting and operating Kubernetes clusters optimised for GPU workloads, leveraging NVIDIA GPU Operator, Network Operator and DCGM * Developing, deploying and maintaining custom Kubernetes operators and controllers to automate infrastructure services * Integrating NVIDIA device plugins, Multi-Instance GPU (MIG) and GPU sharing features into the scheduling layer * Optimising GPU utilisation and job placement through scheduler extensions, such as kube-scheduler plugins, Slurm and Volcano * Collaborating with HPC, ML and DevOps teams to ensure multi-tenant, high-throughput cluster performance * Driving observability and telemetry integrations using Prometheus, Grafana, DCGM Exporter and OpenTelemetry * Implementing secure multi-user and multi-namespace GPU isolation, with RBAC and policy enforcement, such as OPA or Gatekeeper * Maintaining CI/CD pipelines for Kubernetes infrastructure using GitOps, ArgoCD and FluxCD * Contributing to infrastructure-as-code, using Terraform, Helm, and Kustomize * Participating in performance tuning, incident response and production readiness reviews ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe)