Fleet Operations Manager, Data Center Infrastructure

The Meta Game, Inc.
Kuna, ID, United States
about 2 months ago
Apply on www.businessworkforce.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$163,000.0 - $238,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Data Analysis Big Data Data Centers Linux Data Analytics Hardware Infrastructure

Job description

Meta is seeking a forward-thinking, experienced individual to join the Data Center Fleet Operations team. The Fleet Operations Manager is accountable for managing and leading a geographically dispersed team, delivering SLA/KPI’s related to production server hardware, resolution of systemic technical issues, and repairs throughout the assigned geographic region of data centers. We are looking for someone who can effectively prioritize and adapt to shifting priorities in a dynamic operational environment. The ideal candidate is an IT professional with strong leadership skills and experience in Server Hardware, Project Management, Quality Management, Data Analytics, Networks, OS repair, Linux and Automation, ideally in a datacenter environment. Having an extensive understanding of managing servers in a large-scale distributed environment, 1. Build and lead a geographically dispersed, high-performing data center operations team, developing both the technical capabilities and leadership qualities of engineers

  1. Establish and manage a Data Center Operations Team accountable for the maintenance and operation of server hardware and supporting infrastructure at scale
  2. Become a technical expert in Meta’s infrastructure, including platforms, tools, systems, architecture, workflows, and performance
  3. Provide strategic direction, guidance, and support for site and fleet-level operations
  4. Analyze and drive continuous improvement in the engineering and operational performance of our data centers
  5. Employ data analytics to identify inefficiencies, opportunities, exceptions, and correlations in a complex, highly interconnected, technical environment. Enable rapid and effective problem solving, along with proactive identification and mitigation of risks and issues
  6. Collaborate with cross-functional partner teams to ensure fleet health and maintain targeted capacity levels, resulting in optimized operations, minimized downtime, and seamless scalability
  7. Evolve and optimize processes in a globally consistent way to allow Meta to scale and grow effectively
  8. Support and mentor engineers in their day-to-day work, as well as in finding opportunities to develop and grow based on their areas of strength and interest
  9. Create and drive a culture of ownership, innovation, collaboration, accountability, continuous improvement, and safety
  10. Conduct performance management for a technical engineering team, providing clear expectations and goals
  11. Assume the role of incident manager during large-scale, site-wide, and region-wide production-impacting events, as the primary point of contact for your site. This requires working cross-functionally to scope problems, mitigate risks, affect fixes, and communicate the nature, status, and resolution plan for incidents
  12. Support and contribute thought leadership to the development and implementation of business practices, processes and automated tooling
  13. Develop deep knowledge and ownership of a hyper-scale computing fleet through the use of data analysis to identify trends and systemic issues and opportunities
  14. reporting out globally and sharing with peers as appropriate

Requirements

  1. BS, BA, or BEng in a technical field or commensurate experience
  2. Ability to travel up to 30% is required
  3. Experience participating in or leading technical projects related to areas such as process improvement, technology, and/or automation, including bringing in additional expertise as needed
  4. 5+ years of experience managing teams of technical resources, including people and performance management responsibilities
  5. Understanding of data center infrastructure and/or operations, including power, cooling, and/or network systems
  6. structured cabling
  7. and management of projects, incidents, and vendors
  8. Experience using data and metrics to drive decision-making
  9. Ability to influence effectively, working on cross-functional teams to advance the needs of the company and adapting teams to meet these needs
  10. 10+ years of engineering or operations experience, preferably in a mature engineering or operations environment, working with cross-functional teams
  11. Ability to communicate effectively, in a clear and concise manner, appropriately tailoring messages to the audience, 1. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  12. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  13. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  14. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  15. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  16. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  17. Six Sigma knowledge/certification
  18. Experience leading technical resources using Linux or an equivalent OS to support hardware systems in a complex IT environment
  19. Experience with large-scale AI implementations and the use of AI to drive automation
  20. Experience in large-scale data center hardware deployments and building scalable infrastructure
  21. Knowledge of the interdependencies of data center functions and technologies, including electrical, cooling, structured cabling, security, and network

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.businessworkforce.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all