Kubernetes Platform AI Infrastructure
Noblesoft Technologies
Santa Clara, CA, United States
9 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Bash Shell
Computer Networks
Data Centers
Software Debugging
Linux
Distributed Computing Environment
Domain Name System (DNS)
Python (Programming Language)
Linux System Administration
Network Control
Node.Js
+10 more
Prometheus
Runbook
Service Discovery
AI Infrastructure
Autoscaling
Grafana
Data Center Networking
AI Platforms
Kubernetes
Bare Metal
Job description
The Candidate will provide senior Kubernetes platform engineering services for AI infrastructure environments supporting model development, distributed training, inference services, and shared platform operations. The role requires strong cluster troubleshooting ability plus pragmatic platform engineering skills in mixed bare-metal and data center environments.
WHAT THIS CANDIDATE WILL BE DOING
- Build, administer, and troubleshoot Kubernetes platforms used for AI and data-intensive workloads.
- Diagnose failures across control plane components, Kubernetes, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtime, and resource isolation.
- Support GPU-enabled Kubernetes environments, including device plugin behavior, driver dependencies, node health, and workload placement.
- Improve platform reliability through automation, standardized configuration, upgrade planning, and cluster validation gates.
- Investigate workload issues involving storage throughput, network policy, DNS, image pulls, autoscaling, pod eviction, and degraded node states.
- Partner with Linux, network, validation, and SRE teams to resolve complex cross-layer failures affecting AI services.
- Create reusable operational runbooks, dashboards, and health checks for day-2 support.
- Contribute to platform hardening, tenant readiness, and service-level objectives.
Requirements
- 7+ years in infrastructure engineering, with deep hands-on Kubernetes administration experience.
- Strong operational understanding of Kubernetes internals and cluster troubleshooting.
- Experience with container runtimes, Helm, GitOps or declarative operations, and cluster lifecycle management.
- Experience supporting GPU workloads on Kubernetes in lab, validation, or production settings.
-
Strong Linux administration foundation and understanding of data center network dependencies.
- Ability to debug issues from symptom to root cause across node, pod, network, storage, and control plane layers.
- Scripting and automation skill in Python, Bash, or Go.
PREFERRED EXPERIENCE
- Experience with Kubeflow, Argo, Prometheus, Grafana, Loki, or service mesh technologies.
- Familiarity with bare-metal Kubernetes and high-performance storage integration.
- Exposure to regulated or high-change-control production environments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
over 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
24 days ago
CS
Christina Schaireiter
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
3 months ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago
BB
Benedikt Bischof
MLOps – What’s the deal behind it?
almost 4 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago