> Markdown version of [/jobs/ext/730487-sr-sre-platform-architect](https://www.wearedevelopers.com/jobs/ext/730487-sr-sre-platform-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. SRE Platform Architect - **Company:** Bitdeer Technologies Group - **Location:** San Jose, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Intelligent Platform Management Interface, BIOS, Cloud Computing, Databases, Data Centers, Firmware, InfiniBand, Subnetting, NetApp Applications, Raw Data, Software Vulnerability Management, AI Platforms, Kubernetes, Information Technology, Slurm, Serverless Computing, Nvme - **Published:** June 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=6f0903526fc5433e ## About the Role Do you have experience in Kubernetes?, Do you have a Bachelor's degree?, * 10+years of production SRE / platform-engineering / infra-architecture, including 3 years at architect level. * Hands-on with GPU / AI-compute infrastructure - NVIDIA GPU ops (DCGM, MIG, vGPU, NVLink/NVSwitch, XID semantics, NCCL), InfiniBand or RoCE fabrics (subnet manager, fabric partitioning, optical health), HPC storage (Lustre, NetApp/Pure/DDN/VAST, NVMe-oF). * Multi-region observability at scale - metrics / logs / traces / profiles / analytics-lake substrate; recording rules, MWMBR burn-rate alerting, SLI/SLO discipline. ## Description Bitdeer is seeking a visionary and hands-on Cloud SRE Architect to lead the design, development, and evolution of our next-generation public cloud platform. This role will oversee the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and AI services, ensuring global scalability, reliability, and performance. The ideal candidate is a strategic thinker with deep technical expertise in cloud infrastructure, platform engineering and AI systems, capable of bridging architecture vision with real-world engineering execution. You will collaborate closely with cross-functional teams and global partners to define our cloud technology roadmap, optimize multi-region deployments, and deliver world-class infrastructure and platform solutions that power large-scale AI and enterprise workloads., Own the end-to-end architecture of the NeoCloud SRE platform - the substrate that observes, protects, and operates a multi-region GPU rental fleet across self-built and OEM-rented data centers. You are the single point of architectural accountability across the platform's ~57 bounded contexts, ~12 frameworks, and three operational tiers (Edge DC Regional Controller Global Hub). This role is for someone who writes the design, defends it under review, and shepherds it through the engineering squads that build it. What You'll Do * Write and maintain the platform architecture document - keep the design coherent across all sections, frameworks, and tiers. The current document is your starting point. * Review every framework-level change - new bounded context, new plugin kind, tier-deployment shift, schema change, naming change, cross-context contract change. Architecture changes ride GitOps PRs like any other artifact. * Set design invariants - residency rules (raw data stays in Region), Tier 2 self-sufficiency budget ( 24 h), survival-uplink contracts, naming conventions, SLO catalogues, redaction-at-boundary rules. * Run the plugin framework - every extension uses one uniform contract (Common + Domain manifest, lifecycle, observability). You author and evolve this contract. * Decide tier placement - what runs at Edge DC vs Regional Controller vs Global Hub, with data-residency / compliance / availability tradeoffs explicit. * Coordinate with cloud-service teams and tenants - they author plugins, SDKs, dashboards, agent recipes that ride the platform. You set the contracts they consume. * Coordinate with Security - joint ownership of vulnerability management, exposure management, joint operations. Security owns policy and risk acceptance; you own the operational mechanisms they ride. * Pre-flight roadmap items - for any new capability, produce a one-page design that fits the existing layered model (L1-L6), tier topology, naming conventions, and extension contracts before implementation starts. * Defend the design under review - say no to scope creep, special-case workarounds, and one-off integrations that don't fit the framework model. Say yes when a new plugin kind is genuinely needed., * Cluster platforms - first-hand experience with Kubernetes (control plane + GPU Operator + topology-aware scheduling) AND at least one of Slurm / Volcano / Kueue / Ray / KubeRay. * Data-center operations - ZTP, BMC/IPMI/Redfish, BIOS/firmware lifecycle, RMA, multi-vendor OEM management (self-built + leased DC mix). * Strong DDD instincts - bounded contexts, public contracts, no shared databases, one-context-one-repo discipline. * Plugin framework design - you have built (or substantively contributed to) a real extension framework with a uniform manifest + lifecycle. * Writing fluency - you can author and maintain a multi-thousand-line architecture document under review without it drifting; you can also write a one-pager an executive will read. * Cross-team operating tempo - design reviews, runbook authorship, on-call shadowing, post-mortem facilitation * Hyperscale or NeoCloud experience * BS/MS in Computer Science or similar ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [How to Submit CFPs and Get into Public Speaking - Moran Weber](https://www.wearedevelopers.com/videos/2112-how-to-submit-cfps-and-get-into-public-speaking-moran-weber) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?](https://www.wearedevelopers.com/magazine/701-what-makes-wearedevelopers-world-congress-different-from-every-other-tech-event)