> Markdown version of [/jobs/ext/2450635-storage-rack-infrastructure-automation-cluster-bring-up-hive-program](https://www.wearedevelopers.com/jobs/ext/2450635-storage-rack-infrastructure-automation-cluster-bring-up-hive-program). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program - **Company:** SanDisk - **Location:** Marvell, AR, United States - **Contract:** Permanent contract - **Skills:** Automation of Tests, Intelligent Platform Management Interface, Dynamic Host Configuration Protocol, Dynamic Random-Access Memory, Ethernet, Apache Hive, Python (Programming Language), Node.Js, Ansible, Ceph (Software), Gitlab-ci, Infrastructure Automation Frameworks, Bare Metal, Terraform, Jenkins - **Published:** August 29, 2026 - **Apply:** https://www.careerjet.com/jobad/usca7c9fb9287062eb28e43a916e3bf057 ## About the Role Core (must-have): * Deep experience in lab hardware discovery, inventory, and automated provisioning at fleet scale. * Bare-metal automation: PXE/DHCP boot, BMC out-of-band management (Redfish/IPMI), image/OS deployment, cloud-init/first-boot. * Infrastructure-as-code and config automation (e.g., Ansible/Terraform-class tooling) with a declarative, reconcile-to-desired-state mindset. * Strong scripting/automation (Python and shell) and CI systems (Jenkins/GitLab CI or equivalent). * Comfort designing systems that are autonomous by default with reliable manual override. Strongly preferred (or ramp-up expected): * Ceph operational knowledge: OSD/MON/MGR/MDS roles, the orchestrator/cephadm model, placement specs, ceph-volume, CRUSH maps and failure domains, MON quorum, and cluster health/lifecycle. * Distributed-systems fluency: quorum/Paxos intuition, rebalance/recovery behavior, failure-domain reasoning. * Networking bring-up for storage fabrics (RoCEv2/Ethernet, port discovery). * Familiarity with DPU/SoC-based nodes and constrained-node environments. ## Description Hive is a swarm of hundreds of identical storage nodes (Marvell/XSight DPU + SSD), each running a full Ceph plane: OSD, Monitor (MON), Manager (MGR), MDS, plus SeaStore and the NFS front end. Standing up, re-configuring, and recovering a cluster of this size by hand does not scale. We need an engineer who owns the automated hardware discovery and Ceph role-assignment pipeline: the flow that inventories every node and its hardware, decides which daemons each node should run, and drives the cluster from bare metal to a serving state. The flow must run fully autonomously by default, with a clean manual-override path for lab, bring-up, and failure-injection scenarios. This is a specialized automation and lab-hardware-discovery discipline that the Hive team does not currently have dedicated ownership for. It is directly on the critical path for every test cluster, every silicon bring-up, and every customer-shaped deployment. What you'll own * Automated hardware discovery. Detect nodes as they power on and enumerate their hardware (DPU/SoC model, SSD media, DRAM, RNIC/network ports, BMC) using out-of-band and in-band inventory (BMC/Redfish/IPMI, PXE/DHCP boot, cloud-init/first-boot agents). Produce a single authoritative machine inventory that the rest of the pipeline consumes. * Role assignment and cluster composition. Given the discovered inventory, decide and apply which Ceph roles each node runs - OSD, MON, MGR, MDS (and the NFS gateway) - encoded as declarative placement specifications. Own MON quorum sizing and placement, MGR redundancy, MDS/gateway placement, and CRUSH map / failure-domain layout so data and metadata land correctly across the swarm. * Bare-metal * serving bring-up. Drive the end-to-end sequence: node provisioning, OS/image deploy, cluster bootstrap, daemon deployment via the Ceph orchestrator (cephadm-style), ceph-volume-style OSD provisioning on the SSD media, and health convergence to HEALTH_OK. * Autonomous and manual modes. Make the default path zero-touch (a rack powers on and self-assembles into a healthy cluster), while exposing deterministic manual controls to pin roles, hold a node out, force a specific topology, or reproduce a customer/lab configuration for testing. * Lifecycle and recovery automation. Node add/remove, drain and rebalance, daemon replacement, MON re-quorum after loss, MDS/OSD failover validation, and re-discovery after re-imaging - integrated with Hive's Fast Recovery + BMC work. * Reconciliation and drift control. Continuously compare declared desired state against observed cluster state and converge - the same idea Ceph's orchestrator applies to service specs - with clear reporting when reality diverges from intent. * CI/lab integration. Wire the pipeline into automated test so any commit can spin up a correctly-composed multi-node cluster on real hardware and tear it down cleanly., Sandisk thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution. Sandisk is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying. ## Related Videos - [From Cloud Racks to Control Cabinets: Operating Kubernetes on Edge Devices](https://www.wearedevelopers.com/videos/100160-from-cloud-racks-to-control-cabinets-operating-kubernetes-on-edge-devices) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Introducing Redis Agent Memory Server](https://www.wearedevelopers.com/magazine/699-introducing-redis-agent-memory-server)