Data Center Capacity Manager

The Meta Game, Inc.
New Albany, OH, United States
7 days ago
Apply on www.businessworkforce.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$160,000.0 - $223,000.0
Working hours
Regular working hours

Tech stack

Adobe InDesign Artificial Intelligence Systems Engineering Bash Shell Big Data Computer Engineering Data Centers Perl (Programming Language) Fault Tolerance Python (Programming Language) Linux System Administration Network Architecture
+5 more
Network Service Runbook AI Infrastructure System Availability Information Technology

Job description

Meta is looking for a forward-thinking people manager to manage rack capacity planning and delivery, to join the Edge and Network Services (ENS) Foundation organization. In this role, you will lead engineers to improve efficiency, reliability, and risk management via process, systems, and data in some of the largest scale data centers in the world. The right candidate will thrive in a fast-moving organization and enjoy digging into complex operational and reliability challenges in order to implement process and technical system solutions at a global scale.This is a unique opportunity to build and scale a global team. You will develop talent, define operational guardrails, and shape a team centered around ownership, precision, and continuous improvement. In this role, you will own and shape the operations of Meta’s next-generation compute infrastructure. This role has direct influence over how we design and deploy the environments that connect our global infrastructure. You will own high-impact scope with visible outcomes. This role blends strategic and hands-on technical leadership. You will drive long-term process and tooling direction while also being close enough to the ground to unblock restoration issues, optimize workflows, and ensure operational performance expectations. This role interfaces with a broad cross functional group of stakeholders, providing the opportunity to influence standards, champion automation, and extend best practices across the subsea ecosystem.The Manager of the Capacity Team is accountable for leading the team through rack receipts, moves, and decommissioning at our rapidly expanding global data centers. The Capacity team represents the outcome of the central capacity planning at each Data Center Site and interacts with all local cross-functional partners, central tooling and planning teams to deliver capacity for Meta’s applications and AI infrastructure., 1. Provide oversight of the capacity function across multiple sites in a geographical area

  1. Distribute the capacity workload based on projections, OKRs, and KPIs across skill-sets, people resources, and company priorities
  2. Work with FieldOps leaders and engineers to ensure on-time delivery with partners of rack SLAs, ensuring throughput objectives are met
  3. Manage stakeholder communications and represent the team in XFN leadership meetings
  4. Ensure standard Methods of Procedure (MOPs), Runbooks, standard procedures and practices are consistently followed to achieve successful deployment metrics
  5. Effectively cascade strategic organizational goals to ensure alignment for team and individual goals
  6. Cultivate a collaborative and high-performing team and ensure accountability for successful new employee on-boarding
  7. Build and sustain an effective global organization that can scale for the future and lead by example encouraging team collaboration around the world
  8. Manage team of engineers responsible for ensuring stable compliant delivery of capacity, including monitoring Key Performance Indicators/Service Level Indicators, triaging incidents, and driving root-cause analysis
  9. Sustain a proactive highly action oriented team via ongoing engagement around team member responsibilities, priorities, performance expectations, and development planning to optimize an individual’s performance
  10. Attract top talent and fill gaps in the existing team quickly to accommodate growth
  11. Structure teams and working groups to optimize for delivery of business outcomes, set clear expectations, establish operational guardrails, and build a team emphasizing ownership, collaboration, and continuous improvement
  12. Create, maintain and enforce Standard Operating Procedures, runbooks, provisioning workflows, and acceptance criteria ensuring all processes adhere to Meta-wide best practices in design, security/compliance, and reliability
  13. Support internal customers (Production Operations, Network Engineering, Systems Engineering, Logistics, Program Management) and vendors (field engineers, logistics, facilities, and hardware partners), ensuring that capacity delivery meets their needs proactively across global sites
  14. Own the planning and execution of the organizational-level roadmap and strategy to deliver business outcomes
  15. Provide root cause analysis and corrective action leadership to resolve issues across infrastructure, architectures and hardware platforms
  16. Represent the organization and manage interaction with third parties such as hardware, facilities vendors, logistics and managed service partners
  17. Lead highly cross-functional infrastructure projects and programs in a matrix organization covering a range of areas (data center, production network, infrastructure, logistics, supply chain, compliance, legal, and software system engineering)
  18. Communicate cross-functionally across various teams, organizations and internal and external stakeholders, at all levels of the organization
  19. Initiate and lead strategic initiatives driving efficiency, innovation, and standardization of network infrastructure within the data center environment
  20. Drive improvements in technical references/standards, design processes, and training documentation
  21. Support automation and tooling initiatives to drive consistency and efficiency in all aspects of infrastructure design, deployment, and operations
  22. 20% domestic and international travel required to data centers managed

Requirements

  1. Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical discipline, or equivalent practical experience
  2. 5+ years of direct experience managing engineers and organizations
  3. 7+ years of work experience with designing and deploying large-scale data center connectivity infrastructure
  4. Thorough knowledge of structured cabling and infrastructure, including design and installation best practices, industry standards, application, and limitations
  5. Working knowledge of fault tolerant critical infrastructure used for supporting network and compute hardware in a high availability environment
  6. General knowledge of the data center design, construction, and start up/commissioning processes
  7. Experience leading cross-functional teams through partnership and driving large-scale data center or similar infrastructure design initiatives
  8. Track record of solving problems, setting strategic direction, executing tactically, and achieving measurable outcomes
  9. Demonstrated experience to work effectively in an ambiguous environment and collaborate with global cross-functional teams to develop and execute plans that address evolving business needs
  10. Proficiency with workload scoping, prioritization, delegation and multitasking
  11. Demonstrated skills in written and verbal communication
  12. ability to successfully present concepts to variety of stakeholders
  13. 7+ years of direct experience understanding and influencing network and server architectures to include constraint and dependency analysis and translating these into deployable and supportable solution requirements
  14. Experience successfully collaborating across a global team and with cross-functional partners (e.g. physical infrastructure, engineering, strategy, security, logistics) at all levels
  15. Experience managing from the front to prioritize and drive the greater mission forward by translating strategy into results
  16. Direct experience in operational compliance, environmental health and safety, physical and logical infrastructure security, and/or business continuity disciplines
  17. 10+ years of direct experience working within global network or infrastructure operations, deployment, design/engineering and/or support teams
  18. 5+ years of direct experience working within varied infrastructure environments, such as data centers, colocation facilities, or corporate campuses
  19. 5+ years of experience partnering with hardware, ITAD, logistics, facility, security and colocation vendors
  20. Experience identifying and managing key metrics, such as SLA/KPIs, to evaluate success and validate the business impact of capacity services, 1. Experience with scripting and automation (Bash, Python, Perl)
  21. Experience working with on-site construction teams, and managing contractors, and vendor relationships
  22. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  23. Experience with ODM hardware and Linux-based systems
  24. Experience leading and managing matrixed technical teams
  25. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  26. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  27. Experience with purchasing, negotiating and end-to-end supplier management, such as managing global Requests For Proposal and contract negotiations
  28. Experience influencing operational requirements in consortia and contract negotiations

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.businessworkforce.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:52 min

Avoiding remote code execution from unsanitized inputs

Alexander Pirker · World Congress 2022

1:29 min

Tech infrastructure capacity and AI product innovations

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · World Congress 2026 Europe

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all