> Markdown version of [/jobs/ext/2631704-kubernetes-platform-ai-infrastructure](https://www.wearedevelopers.com/jobs/ext/2631704-kubernetes-platform-ai-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Kubernetes Platform AI Infrastructure - **Company:** Noblesoft Technologies - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Computer Networks, Data Centers, Software Debugging, Linux, Distributed Computing Environment, Domain Name System (DNS), Python (Programming Language), Linux System Administration, Network Control, Node.Js, Prometheus, Runbook, Service Discovery, AI Infrastructure, Autoscaling, Grafana, Data Center Networking, AI Platforms, Kubernetes, Bare Metal - **Published:** August 27, 2026 - **Apply:** https://www.dice.com/job-detail/bc53a3b8-d5a0-4a4b-83ad-34ac2cc4ff46 ## About the Role * 7+ years in infrastructure engineering, with deep hands-on Kubernetes administration experience. * Strong operational understanding of Kubernetes internals and cluster troubleshooting. * Experience with container runtimes, Helm, GitOps or declarative operations, and cluster lifecycle management. * Experience supporting GPU workloads on Kubernetes in lab, validation, or production settings. * Strong Linux administration foundation and understanding of data center network dependencies. * Ability to debug issues from symptom to root cause across node, pod, network, storage, and control plane layers. * Scripting and automation skill in Python, Bash, or Go. PREFERRED EXPERIENCE * Experience with Kubeflow, Argo, Prometheus, Grafana, Loki, or service mesh technologies. * Familiarity with bare-metal Kubernetes and high-performance storage integration. * Exposure to regulated or high-change-control production environments. ## Description The Candidate will provide senior Kubernetes platform engineering services for AI infrastructure environments supporting model development, distributed training, inference services, and shared platform operations. The role requires strong cluster troubleshooting ability plus pragmatic platform engineering skills in mixed bare-metal and data center environments. WHAT THIS CANDIDATE WILL BE DOING * Build, administer, and troubleshoot Kubernetes platforms used for AI and data-intensive workloads. * Diagnose failures across control plane components, Kubernetes, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtime, and resource isolation. * Support GPU-enabled Kubernetes environments, including device plugin behavior, driver dependencies, node health, and workload placement. * Improve platform reliability through automation, standardized configuration, upgrade planning, and cluster validation gates. * Investigate workload issues involving storage throughput, network policy, DNS, image pulls, autoscaling, pod eviction, and degraded node states. * Partner with Linux, network, validation, and SRE teams to resolve complex cross-layer failures affecting AI services. * Create reusable operational runbooks, dashboards, and health checks for day-2 support. * Contribute to platform hardening, tenant readiness, and service-level objectives. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From Factory Floor to Kubernetes Core: Building an Edge Platform One Step at a Time](https://www.wearedevelopers.com/videos/1415-from-factory-floor-to-kubernetes-core-building-an-edge-platform-one-step-at-a-time) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)