> Markdown version of [/jobs/ext/2272197-senior-kubernetes-platform-engineer-ai-infrastructure](https://www.wearedevelopers.com/jobs/ext/2272197-senior-kubernetes-platform-engineer-ai-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Kubernetes Platform Engineer - AI Infrastructure - **Company:** STRYDE CONSULTING SERVICES LLC - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Computer Networks, Data Centers, Linux, Distributed Computing Environment, Domain Name System (DNS), Python (Programming Language), Linux System Administration, Node.Js, Octopus Deploy, Prometheus, Azure Machine Learning, Service Discovery, AI Infrastructure, Graphics Processing Unit (GPU), Istio, System Availability, Grafana, Data Center Networking, Kubernetes, Bare Metal, Hardware Infrastructure - **Published:** August 27, 2026 - **Apply:** https://www.dice.com/job-detail/c5e94909-ada5-478a-ab65-289d37267846 ## About the Role * 7+ years of infrastructure/platform engineering experience with strong hands-on Kubernetes administration. * Deep understanding of Kubernetes architecture and internals. * Strong experience troubleshooting Kubernetes clusters and production workloads. * Hands-on experience with: + Kubernetes + Container runtimes + Helm + CNI and CSI + Ingress + Kubernetes scheduling + Node lifecycle management + Cluster upgrades and lifecycle management * Experience with GPU workloads on Kubernetes in production, validation, lab, or AI infrastructure environments. * Strong Linux administration and troubleshooting skills. * Understanding of data center networking and storage dependencies. * Ability to troubleshoot issues from initial symptoms through root cause across multiple infrastructure layers. * Experience with declarative infrastructure, GitOps, or configuration-driven operations. * Strong scripting and automation skills using Python, Bash, or Go. * Strong communication and collaboration skills with infrastructure and engineering teams. * Ability to work independently in high-severity and ambiguous production situations. Preferred Skills * Kubeflow * Argo / Argo CD * Prometheus * Grafana * Loki * Service mesh technologies * Bare-metal Kubernetes * High-performance storage * GPU cluster infrastructure * NVIDIA GPU ecosystem * AI/ML platform infrastructure * Distributed training environments * Regulated or high-change-control production environments Ideal Candidate The ideal candidate is a hands-on Kubernetes platform engineer who can operate beyond standard cluster administration and troubleshoot complex AI infrastructure issues involving Kubernetes, Linux, GPUs, networking, storage, and compute. Candidates should be comfortable working directly with bare-metal infrastructure and data center environments and collaborating across platform, Linux, networking, storage, validation, and SRE teams. ## Description We are seeking a Senior Kubernetes Platform Engineer to support AI infrastructure environments focused on model development, distributed training, inference services, and shared platform operations. The ideal candidate will have deep hands-on Kubernetes administration and troubleshooting experience, strong Linux fundamentals, and experience supporting GPU-enabled workloads in bare-metal and data center environments. This is a hands-on platform engineering role requiring the ability to troubleshoot complex issues across Kubernetes, compute, networking, storage, containers, and GPU infrastructure., * Build, administer, maintain, and troubleshoot Kubernetes platforms supporting AI and data-intensive workloads. * Troubleshoot Kubernetes control plane components, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtimes, and resource isolation. * Support GPU-enabled Kubernetes environments, including device plugins, GPU drivers, node health, resource allocation, and workload placement. * Troubleshoot cluster and workload issues involving storage performance, networking, DNS, network policies, image pulls, autoscaling, pod evictions, and degraded nodes. * Perform root-cause analysis across nodes, pods, networking, storage, GPU infrastructure, and Kubernetes control plane components. * Improve platform reliability through automation, standardized configurations, upgrade planning, and cluster validation. * Manage Kubernetes cluster lifecycle activities, including upgrades, configuration changes, health checks, and validation. * Partner with Linux, network, storage, validation, and SRE teams to resolve complex cross-layer infrastructure issues. * Develop reusable operational runbooks, dashboards, monitoring, and health checks for day-2 operations. * Contribute to platform hardening, tenant readiness, reliability, and service-level objectives. * Automate platform administration and operational tasks using Python, Bash, or Go. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Stop Using Node.js Like It’s 2020! - Alfonso Graziano](https://www.wearedevelopers.com/videos/1863-stop-using-node-js-like-it-s-2020-alfonso-graziano) ## Related Articles - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)