> Markdown version of [/jobs/ext/2073233-k8-site-reliability-sme](https://www.wearedevelopers.com/jobs/ext/2073233-k8-site-reliability-sme). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # K8 Site Reliability SME - **Company:** Bitdeer Technologies Group - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Clusters, Computer Programming, Computer Networks, Python (Programming Language), Role-Based Access Control, Reliability Engineering, Prometheus, Runbook, Graphics Processing Unit (GPU), Grafana, Backend, Kubernetes, Bare Metal, Slurm, Terraform, Pagerduty - **Published:** August 15, 2026 - **Apply:** https://www.wayup.com/i-j-K8-Site-Reliability-SME-Bitdeer-Technologies-Group-244525614950925/ ## About the Role + 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S + Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S + Experience with topology-aware scheduling and GPU-specific resource management + Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees + Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom) + Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux) + Strong SRE background: SLI/SLO frameworks, incident management, capacity planning + Experience with Prometheus, Grafana, and alerting at scale + Strong programming skills in Go or Python for operator/CRD development + AIOps aptitude - you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one. + Runbook-as-code mindset - every SRE playbook you write should be executable by the platform. ## Description You run the control plane where AIOps meets tenants - where topology-aware scheduling, self-healing, and agent-driven remediation actually execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane - and you make sure the AIOps substrate can reach in and remediate without a human on the pager. What you'll own + Production Kubernetes clusters optimized for GPU workloads at scale (100-10,000 GPUs). + Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies. + Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity. + Custom Resource Definitions (CRDs) for GPU workload lifecycle management. + AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow. + Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards. + Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation. + Terraform providers and modules for infrastructure-as-code across GPU clusters. + SLIs/SLOs for cluster availability, job completion rates, and provisioning latency. + Incident management: runbook automation, escalation, post-incident reviews. + Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty. + GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling. Feed the AIOps substrate + The remediation-actuator and workflow engine land here - you make the control plane safe for automated action. + Your CRDs are the schema the platform's predictors and remediators write against. + Every human intervention you do this quarter becomes an autonomous workflow next quarter. What success looks like in year 1 + Automated drain/reschedule around predicted GPU faults, at scale, without customer impact. + BMaaS live for external tenants with self-service onboarding. + Cluster availability and job-completion SLOs published and met. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)