> Markdown version of [/jobs/ext/2068252-ai-storage-solutions-expert](https://www.wearedevelopers.com/jobs/ext/2068252-ai-storage-solutions-expert). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Storage Solutions Expert - **Company:** Bitdeer Technologies Group - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Transmissions, Data Migration, Linux, Distributed Data Store, Firmware, Storage Area Network (SAN), Performance Tuning, Remote Direct Memory Access, Runbook, Weka, Ceph (Software), Mttr, Storage Technologies, Nvme - **Published:** August 15, 2026 - **Apply:** https://www.wayup.com/i-j-AI-Storage-Solutions-Expert-Bitdeer-Technologies-Group-450465788567424/ ## About the Role + 5+ years in enterprise or HPC storage operations, with at least 2 years supporting AI/ML workloads + Hands-on deployment and operations experience with at least two of: WEKA, VAST Data, Ceph, DDN/Lustre + Strong understanding of AI training I/O patterns: checkpoint frequency, dataset loading, shuffle buffers + Experience with high-performance storage networking (NFS over RDMA, NVMe-oF) + Knowledge of GPU Direct Storage and RDMA-based data transfer + Proficiency in storage performance benchmarking and tuning (fio, IOR, mdtest) + Experience implementing multi-tenant storage with isolation and QoS + Strong Linux systems knowledge (kernel tuning, filesystem internals, block device management) + Instinct for turning ops toil into ML signal - you've either shipped an anomaly detector for storage/IO telemetry or you can articulate the labels and features you'd need to. + Runbook-as-code mindset - every SOP you write should be executable by a machine within a quarter. ## Description You own the IO layer that trains the models - and the signals we need to predict storage faults before a checkpoint stalls a $50M training run. Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a slow parallel read can starve a 1,000-GPU job; a stalled checkpoint can waste a full training epoch. In this role you deploy and operate the high-performance storage layer for AI training and inference across NeoCloud's US DCs, and you feed the AIOps substrate with the signals it needs to catch storage regressions before they page a customer. What you'll own + Deploy and operate parallel/distributed storage systems: WEKA, VAST Data, Ceph, DDN/Lustre. Design storage architectures optimized for AI workload patterns - checkpoint I/O bursts, sequential dataset reads, KV cache for inference. + Implement multi-tenant storage isolation with per-tenant QoS, quotas, and access controls; configure and optimize GPU Direct Storage for direct GPU-to-storage data paths. + Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed storage fabrics) and Nvidia CMX for cluster-wide storage orchestration. + Diagnose and tune storage performance: IOPS, throughput, latency profiling with fio, IOR, mdtest; own the runbook for common failure modes. + Plan storage capacity aligned with GPU cluster growth and customer workload projections; manage firmware, data migration, and DR procedures. Feed the AIOps substrate + Instrument storage telemetry - IO tail latency, checkpoint durations, NVMe SMART, filesystem health, RDMA counters - into the metrics/logs/traces store the platform team runs. + Partner with the platform team to define the storage-fault predictor: which signals, which labels (from your incidents), which false-positive tolerances. + Convert every novel incident into an automation: SOPs become runbook-as-code, runbook-as-code becomes an agent-executable remediation. What success looks like in year 1 + Observability and a baseline predictor for the top 3 storage-fault classes on our fabric. + Storage-incident MTTR measurably lower than at hire. + The Nvidia GB200-class clusters we build out ship on your storage design. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [TiDB, One Layer at a Time: How Distributed SQL Became an Agentic AI Backbone](https://www.wearedevelopers.com/videos/100117-tidb-one-layer-at-a-time-how-distributed-sql-became-an-agentic-ai-backbone) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)