> Markdown version of [/jobs/ext/2710057-systems-oriented-cluster-systems-capacity-engineer](https://www.wearedevelopers.com/jobs/ext/2710057-systems-oriented-cluster-systems-capacity-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # systems-oriented Cluster & Systems Capacity Engineer - **Company:** Backblaze, Inc. - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $123,000.0 - $175,000.0 - **Contract:** Permanent contract - **Skills:** Microsoft Excel, Data Analysis, Cloud Computing, Cloud Storage, Information Systems, Databases, Data Centers, Distributed Systems, Python (Programming Language), Network Architecture, Reliability Engineering, Prometheus, SQL Databases, Snowflake, Grafana, Information Technology - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/cluster-systems-capacity-engineer-backblaze-external-website-8612347 ## About the Role * Bachelor's degree in Computer Science, Engineering, Mathematics, Data Science, Information Systems, Statistics or a related, technical field (or equivalent experience). * 3-6+ years of experience in Site Reliability Engineering, Infrastructure Capacity Planning, Systems/Infrastructure Engineering, Production Engineering, Data Center Operations or similar Cloud Operations role * Familiarity and experience working with Cloud Storage infrastructure, particularly highly-available, large-scale distributed systems supporting large amounts of data with high throughput and complex performance requirements * Background in capacity modeling, performance analysis, scenario modeling, and/or infrastructure cost optimization, with an ability to quantify ideas within financial frameworks and forecasts. * Proficiency in database and data analysis tools (preferably Snowflake, Metabase, Grafana, Python, SQL, Prometheus, Victoria Metrics, and Excel/Google Sheets) * Demonstrated deep, creative, and logical thinking complimented by a strong data analysis skillset * Excellent communication and documentation skills, with the ability to share knowledge and explain concepts accurately and concisely * Desire to work on a highly-autonomous team that cares deeply about quality, cost, and the customer experience ## Description This role ensures that Backblaze's storage clusters, compute systems, and network infrastructure scale reliably, cost-efficiently, and ahead of demand. You will build and maintain predictive models, ensure consistent supply and demand alignment, and partner cross-functionally to inform strategic investment and deployment decisions. ., Capacity Planning & Forecasting * Develop and maintain short, medium, and long-term capacity demand and hardware deployment forecasts across storage, compute, and network domains within the platform * Build predictive models that translate business demand signals into infrastructure requirements using historical utilization, growth trends, product sales plans, hardware lifecycle roadmaps, and other key business inputs * Partner with Infrastructure, Production, and Network Engineering teams to align capacity plans with system design and scaling initiatives * Develop and automate forecasting pipelines, simulation calculators and tools, and capacity dashboards to improve data quality, reduce manual analysis, and provide stakeholders clear visibility into platform usage and cluster health metrics Cluster Performance & Resource Optimization * Monitor and analyze cluster and system-level utilization and performance across CPU, memory, IOPS, and network resources * Adjust deployment plans and recommended configurations in real-time to maintain adequate headroom and system stability in support of delivering a world-class customer experience * Partner with service and platform owners to develop headroom and live buffer policies, optimize hardware BoMs, leverage virtualized orchestration, and reduce product cost Cross-functional Organizational Alignment * Work in lockstep with Operations and Finance peers to align capacity plans and hardware requirements with capital budgets, cost targets, and financial outcomes * Support strategic optimization initiatives across infrastructure investments, engineering development, and operations processes, contributing to long-term infrastructure strategy and capital planning * Lead efforts to evaluate, procure, and provision requests for new or additional hardware, working with Systems and Network Engineering, SRE, NOC, and Data Center Operations teams to identify and deliver optimal solutions * Maintain alignment with Product and Sales to support customer onboarding, growth, and demand variability * Communicate complex capacity and infrastructure insights clearly to technical and non-technical stakeholders ## Related Videos - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Hacking AI at the Edge of the Indian Ocean](https://www.wearedevelopers.com/videos/100177-hacking-ai-at-the-edge-of-the-indian-ocean) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [What’s the Difference Between Frontend and Backend Development?](https://www.wearedevelopers.com/magazine/240-what-s-the-difference-between-frontend-and-backend-development)