Research Scientist, Infrastructure Modeling and Reliability

Facebook Inc.
Menlo Park, CA, United States
25 days ago
Apply on www.jobmonkeyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Data Analysis C++ (Programming Language) Data Centers Distributed Systems R (Programming Language) Monitoring of Systems Python (Programming Language) Machine Learning Reliability Engineering

Job description

  • Define the scientific and technical strategy for modeling power consumption, peak risk, and reliability tradeoffs across large-scale infrastructure systems.
  • Develop statistical, machine learning, and/or optimization models that forecast power demand, estimate peak distributions, quantify uncertainty, and support operational decision-making.
  • Build approaches that reason about high-dimensional signals, correlated demand, failure-domain constraints, reserve margins, and reliability targets.
  • Partner with engineering, capacity planning, data center, energy, hardware, operations, and finance teams to translate model outputs into infrastructure planning and utilization decisions.
  • Establish evaluation frameworks, backtesting methods, confidence intervals, and monitoring systems to measure model quality and operational risk.
  • Identify opportunities to safely increase power utilization, reduce stranded capacity, improve cost efficiency, and guide long-term infrastructure investment.
  • Lead ambiguous, company-critical technical initiatives across organizations, influencing strategy and aligning stakeholders around scientifically grounded decisions.
  • Mentor senior scientists and engineers, raise the technical bar for modeling and forecasting systems, and represent Meta’s work through appropriate external publications, talks, or industry engagement.

Requirements

Meta builds technologies that help people connect, find communities, and grow businesses. Meta’s infrastructure supports services used by billions of people, and operating that infrastructure efficiently requires increasingly sophisticated modeling of demand, utilization, reliability, and physical resource constraints.We are seeking an industry-leading Research Scientist or Applied Scientist to define and build new modeling approaches for power utilization across Meta’s infrastructure. This role will lead the development of statistical and machine learning models that monitor power consumption, project peak demand, quantify uncertainty, and inform how Meta maximizes usable power within failure domains while maintaining target reliability levels. The ideal candidate has deep experience modeling high-dimensional, noisy, and interdependent systems, and has demonstrated the ability to translate scientific advances into production systems that influence large-scale infrastructure strategy., * 10+ years of experience developing statistical, machine learning, simulation, forecasting, optimization, or other quantitative modeling systems

  • Experience leading ambiguous, cross-functional technical programs from problem definition through model development, evaluation, deployment, and business impact
  • Experience coding in Python, R, C++, Java, or similar languages for data analysis, modeling, simulation, or production systems
  • Experience communicating complex technical concepts, assumptions, uncertainty, and tradeoffs to technical and non-technical audiences
  • Experience influencing technical strategy across multiple teams or organizations, * Experience modeling high-dimensional, sparse, noisy, or strongly correlated data in production environments
  • Experience with time-series forecasting, probabilistic forecasting, Bayesian modeling, extreme-value modeling, causal inference, stochastic processes, simulation, or uncertainty quantification
  • Experience with infrastructure, capacity planning, power systems, energy systems, data centers, reliability engineering, distributed systems, supply-chain optimization, or resource allocation
  • Experience building models that support operational decisions under explicit reliability, safety, cost, or utilization constraints
  • Experience developing peak-demand forecasts, confidence intervals, risk estimates, anomaly detection, or backtesting frameworks
  • Experience applying optimization, operations research, or decision science to large-scale resource planning
  • Demonstrated record of industry-level technical leadership, such as defining new research directions, influencing company strategy, publishing in leading venues, or shaping external technical standards
  • Experience mentoring senior technical contributors and building scientific communities across organizations

About the company

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today-beyond the constraints of screens, the limits of distance, and even the rules of physics.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

1:48 min

Automating exploratory data analysis within training pipelines

Dora Petrella · World Congress 2023

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · World Congress 2024

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all