> Markdown version of [/jobs/ext/2617679-storage-and-datacenter-team-lead](https://www.wearedevelopers.com/jobs/ext/2617679-storage-and-datacenter-team-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Storage and Datacenter Team Lead - **Company:** The Voleon Group - **Location:** Berkeley, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $215,000.0 - $245,000.0 - **Contract:** Permanent contract - **Skills:** Bash Shell, Big Data, CentOS, Configuration Management, Databases, Continuous Integration, Data Centers, Data Center Infrastructure Management (CIM), Linux, DevOps, RAID, File Systems, Distributed Data Store, VMware ESX Servers, Identity and Access Management, Networking Hardware, Python (Programming Language), Kernel-Based Virtual Machine, Lightweight Directory Access Protocols (LDAP), PostgreSQL, Linux System Administration, Linux Servers, Nagios, Network Attached Storage (Server Appliance), Performance Tuning, Red Hat Enterprise Linux, Ansible, Prometheus, Virtualization Technology, Ceph (Software), CheckMK, Grafana, Database Optimization, Git, Containerization, Infrastructure Automation Frameworks, Storage Technologies, Docker - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=674817f5b3bb3602 ## About the Role The ideal candidate will bring a deep understanding of Linux-based storage systems, excellent problem-solving skills, and a passion for building reliable, scalable infrastructure. You should have proven experience managing Ceph or similar distributed storage systems, as well as handling large-scale data lifecycle processes including archiving and backup. You will be expected to lead by example-driving automation efforts, and contributing to high-level planning while managing team priorities and operations., * 5+ years of Linux Systems Administration experience with significant recent focus on storage systems * 2+ years of team leadership, technical project management, or mentoring experience * Knowledge of distributed storage systems such as Ceph and storage technologies including RAID, SAN, and NAS * Experience streamlining data lifecycle processes, including archiving, backup, and retention of PB-scale data * Hands-on experience with co-located data center infrastructure * Ability to travel to remote datacenter sites when needed * Strong scripting/development experience in Bash and/or Python * Experience with configuration management tools (Ansible) and infrastructure automation * Familiarity with monitoring and alerting systems (Nagios/CheckMK, Prometheus, Grafana) * Understanding of virtualization (KVM, ESXi) and containerization (Docker, Podman) * Knowledge of LDAP/IPA/AD and centralized identity management Preferred Qualifications * Experience with Kubernetes container orchestration * PostgreSQL DBA experience * Experience in a high-throughput research or trading environment * Exposure to RHEL/CentOS/Rocky Linux in enterprise settings * Experience with DCIM tools for tracking assets, power, and space * Familiarity with CI/CD pipelines and DevOps principles * Experience managing colocation vendor relationships and SLAs ## Description We are seeking a hands-on and strategic Storage and Datacenter Team Lead to guide and grow our critical Storage Engineering team. This individual will be both a technical expert and a team leader, providing architectural oversight, mentorship, and direct implementation support., Leadership & Team Management * Lead a small team of storage, database, and systems administrators with duties including mentorship, performance management, and career development * Coordinate datacenter operations across production and research facilities, including scheduling site work and managing vendor/contractor visits * Align team priorities with organizational goals and ensure timely delivery of projects * Participate in hiring efforts to grow and evolve the storage engineering team * Coordinate on-call schedules and ensure effective incident response processes are in place Technical & Operational Oversight * Architect, implement, and maintain highly available and performant storage systems * Define and drive automation strategies for storage deployment and monitoring * Oversee storage lifecycle management including capacity planning, performance tuning, and data protection strategies such as archiving and backups for large-scale datasets * Provide architectural guidance and hands-on support for Ceph at PB scale * Oversee physical datacenter infrastructure including rack layout planning, power distribution, cooling systems, and capacity forecasting for space, power, and cooling * Manage equipment installation and decommissioning * Collaborate closely with networking, virtualization, research, and application teams to support diverse compute and storage needs * Participate in and improve CI/CD and configuration management processes with tools like Ansible and Git * Support database operations through database tuning, storage optimization, and collaboration with developers * Develop and maintain runbooks for remote-hands work and coordinate with contractors and facility personnel to perform onsite operations IC-Level Engineering & Troubleshooting * Serve as an escalation point for advanced troubleshooting of distributed filesystems, databases, and high-performance storage infrastructure * Hands-on administration of Linux servers, network-attached storage, virtualization platforms, and cluster frameworks * Installation, cabling, and troubleshooting of physical server, storage, and network hardware in rack environments * Diagnose and resolve hardware-level issues impacting production systems * Support and enhance observability using tools such as Prometheus, Grafana, and others ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [It's all about the Data](https://www.wearedevelopers.com/videos/425-it-s-all-about-the-data) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Operating etcd for Managed Kubernetes](https://www.wearedevelopers.com/videos/1191-operating-etcd-for-managed-kubernetes) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [What does the history of data storage tell us about the future?](https://www.wearedevelopers.com/magazine/495-what-does-the-history-of-data-storage-tell-us-about-the-future) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)