Kubernetes Platform AI Infrastructure

Noblesoft Technologies
Santa Clara, CA, United States
9 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Bash Shell Computer Networks Data Centers Software Debugging Linux Distributed Computing Environment Domain Name System (DNS) Python (Programming Language) Linux System Administration Network Control Node.Js
+10 more
Prometheus Runbook Service Discovery AI Infrastructure Autoscaling Grafana Data Center Networking AI Platforms Kubernetes Bare Metal

Job description

The Candidate will provide senior Kubernetes platform engineering services for AI infrastructure environments supporting model development, distributed training, inference services, and shared platform operations. The role requires strong cluster troubleshooting ability plus pragmatic platform engineering skills in mixed bare-metal and data center environments.

WHAT THIS CANDIDATE WILL BE DOING

  • Build, administer, and troubleshoot Kubernetes platforms used for AI and data-intensive workloads.
  • Diagnose failures across control plane components, Kubernetes, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtime, and resource isolation.
  • Support GPU-enabled Kubernetes environments, including device plugin behavior, driver dependencies, node health, and workload placement.
  • Improve platform reliability through automation, standardized configuration, upgrade planning, and cluster validation gates.
  • Investigate workload issues involving storage throughput, network policy, DNS, image pulls, autoscaling, pod eviction, and degraded node states.
  • Partner with Linux, network, validation, and SRE teams to resolve complex cross-layer failures affecting AI services.
  • Create reusable operational runbooks, dashboards, and health checks for day-2 support.
  • Contribute to platform hardening, tenant readiness, and service-level objectives.

Requirements

  • 7+ years in infrastructure engineering, with deep hands-on Kubernetes administration experience.
  • Strong operational understanding of Kubernetes internals and cluster troubleshooting.
  • Experience with container runtimes, Helm, GitOps or declarative operations, and cluster lifecycle management.
  • Experience supporting GPU workloads on Kubernetes in lab, validation, or production settings.
  • Strong Linux administration foundation and understanding of data center network dependencies.

  • Ability to debug issues from symptom to root cause across node, pod, network, storage, and control plane layers.
  • Scripting and automation skill in Python, Bash, or Go.

PREFERRED EXPERIENCE

  • Experience with Kubeflow, Argo, Prometheus, Grafana, Loki, or service mesh technologies.
  • Familiarity with bare-metal Kubernetes and high-performance storage integration.
  • Exposure to regulated or high-change-control production environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:30 min

Overlooked AI infrastructure and operational deployment barriers

Stan Girard Stan Girard · World Congress 2024

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all