Software Engineer, Systems & ML Infrastructure - MSL FAIR Foundations

The Meta Game, Inc.
Menlo Park, CA, United States
2 days ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$154,003.0 - $217,000.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence C++ (Programming Language) Cloud Computing Profiling Code Review Software Debugging Programming Tools Distributed Systems Python (Programming Language) Machine Learning Performance Tuning
+10 more
Reliability Engineering Azure Machine Learning Software Engineering System Programming Workflow Management Systems Reinforcement Learning Data Processing Reliability of Systems Backend Machine Learning Operations

Job description

Meta is seeking Software Engineers to join the Frontier Evals Research team within Meta Superintelligence Labs. Evaluations are a critical part of AI progress at Meta Superintelligence Labs, determining what capabilities get built, which features get prioritized, and how quickly our models improve. As a Systems and ML Infrastructure Engineer on this team, you will build the platforms and services that enable reliable evaluation of our most advanced AI models across text, vision, audio, and beyond. You’ll work alongside researchers and engineers to turn rapidly evolving research workflows into scalable, dependable infrastructure.This is a highly technical software engineering role focused on distributed systems, developer infrastructure, and production-grade ML platforms. You will design and own systems for scheduling and executing evaluation workloads, managing datasets and model artifacts, monitoring correctness and performance, and making results reproducible and easy to consume. The infrastructure you build will directly support research decisions and major model lines within MSL, making reliability, scalability, operational excellence, and engineering rigor paramount.You will succeed by moving quickly in an open-ended research environment while building durable systems, reducing operational toil, and creating abstractions that help researchers iterate faster. If you are passionate about building the technical foundation for frontier AI development and thrive in fast-paced, high-impact environments, we encourage you to apply.

Required Skills:

Software Engineer, Systems & ML Infrastructure - MSL FAIR Foundations Responsibilities:

  1. Design, build, and operate scalable infrastructure for running evaluations across large model fleets, datasets, modalities, and compute environments
  2. Develop orchestration, scheduling, data, and artifact-management systems that make the evaluation workflows reliable and reproducible
  3. Build APIs, abstractions, and developer tools that allow researchers to launch, debug, compare, and interpret evaluations efficiently
  4. Improve system reliability through testing, observability, capacity planning, performance optimization, and automated failure recovery
  5. Partner with research and engineering teams to translate new evaluation requirements into reusable platform capabilities

Requirements

  1. 3+ years of software engineering experience in building backend, distributed, data, or machine learning infrastructure
  2. Proficiency in Python, C++, or another systems programming language
  3. Experience designing, implementing, and operating reliable services, platforms, or data-processing systems
  4. Experience independently delivering medium- to large-scale technical projects from design through production operation
  5. Demonstrated knowledge of software engineering practices, including testing, code review, observability, incident response, and performance analysis
  6. Ability to work effectively with researchers and engineers and to adapt to rapidly changing requirements, 1. Experience building infrastructure for large-scale machine learning training, inference, evaluation, or data processing
  7. Experience with distributed compute systems, workflow orchestration, containers, cluster schedulers, or cloud infrastructure
  8. Experience with performance profiling, resource efficiency, reliability engineering, and production observability
  9. Familiarity with language model post-training workflows, including supervised fine-tuning, reinforcement learning, evaluation, and inference, and the infrastructure needed to support them at scale

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · World Congress 2026 Europe

7:10 min

Exploring pathways into the machine learning engineering field

Jose Luis Latorre Millas · LIVE

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · World Congress 2026 Europe

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · World Congress 2024

Videos

See all

Related articles

See all