Kubernetes platform engineer
IT RECRUIT, LLC
Raleigh, NC, United States
8 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Bash Shell
Computer Networks
Data Centers
Software Debugging
Linux
Domain Name System (DNS)
Python (Programming Language)
Linux System Administration
Node.Js
Prometheus
Runbook
+6 more
Service Discovery
Autoscaling
Istio
Grafana
Kubernetes
Bare Metal
Job description
- Build, administer, and troubleshoot Kubernetes platforms supporting AI and data-intensive workloads.
- Diagnose failures across control plane components, kubelet, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtime, and resource isolation.
- Support GPU-enabled Kubernetes environments, including device plugin behavior, driver dependencies, node health, and workload placement.
- Improve reliability through automation, standardized configuration, upgrade planning, and cluster validation gates.
- Investigate issues involving storage throughput, network policy, DNS, image pulls, autoscaling, pod eviction, and degraded node states.
- Partner with Linux, networking, validation, and SRE teams to resolve cross-layer failures.
- Create reusable runbooks, dashboards, and health checks for day-2 support.
- Contribute to platform hardening, tenant readiness, and service-level objectives.
Requirements
- 7+ years of infrastructure engineering experience.
- Deep hands-on Kubernetes administration experience.
- Strong understanding of Kubernetes internals and cluster troubleshooting.
- Experience with container runtimes, Helm, GitOps or declarative operations, and cluster lifecycle management.
- Experience supporting GPU workloads on Kubernetes.
- Strong Linux administration background.
- Understanding of data-center networking dependencies.
- Ability to debug from symptom to root cause across node, pod, network, storage, and control-plane layers.
- Scripting and automation skills in Python, Bash, or Go.
Preferred
- Kubeflow
- Argo
- Prometheus
- Grafana
- Loki
- Service mesh technologies
- Bare-metal Kubernetes
- High-performance storage integration
- Regulated or high-change-control production environments
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
over 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
24 days ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
CS
Christina Schaireiter
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
3 months ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
CH
Chris Heilmann
Dev Digest 121 - AI goes offline
over 2 years ago