Kubernetes platform engineer

IT RECRUIT, LLC
Raleigh, NC, United States
8 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Bash Shell Computer Networks Data Centers Software Debugging Linux Domain Name System (DNS) Python (Programming Language) Linux System Administration Node.Js Prometheus Runbook
+6 more
Service Discovery Autoscaling Istio Grafana Kubernetes Bare Metal

Job description

  • Build, administer, and troubleshoot Kubernetes platforms supporting AI and data-intensive workloads.
  • Diagnose failures across control plane components, kubelet, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtime, and resource isolation.
  • Support GPU-enabled Kubernetes environments, including device plugin behavior, driver dependencies, node health, and workload placement.
  • Improve reliability through automation, standardized configuration, upgrade planning, and cluster validation gates.
  • Investigate issues involving storage throughput, network policy, DNS, image pulls, autoscaling, pod eviction, and degraded node states.
  • Partner with Linux, networking, validation, and SRE teams to resolve cross-layer failures.
  • Create reusable runbooks, dashboards, and health checks for day-2 support.
  • Contribute to platform hardening, tenant readiness, and service-level objectives.

Requirements

  • 7+ years of infrastructure engineering experience.
  • Deep hands-on Kubernetes administration experience.
  • Strong understanding of Kubernetes internals and cluster troubleshooting.
  • Experience with container runtimes, Helm, GitOps or declarative operations, and cluster lifecycle management.
  • Experience supporting GPU workloads on Kubernetes.
  • Strong Linux administration background.
  • Understanding of data-center networking dependencies.
  • Ability to debug from symptom to root cause across node, pod, network, storage, and control-plane layers.
  • Scripting and automation skills in Python, Bash, or Go.

Preferred

  • Kubeflow
  • Argo
  • Prometheus
  • Grafana
  • Loki
  • Service mesh technologies
  • Bare-metal Kubernetes
  • High-performance storage integration
  • Regulated or high-change-control production environments

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:54 min

Speaker background and open source Kubernetes edge computing projects

Gaurav Gahlot Gaurav Gahlot · World Congress 2026 Europe

7:15 min

Installing Istio programmatically with bash scripts

Thomas Südbröcker · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all