Software Engineer, Systems & ML Infrastructure - MSL FAIR Foundations
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Meta is seeking Software Engineers to join the Frontier Evals Research team within Meta Superintelligence Labs. Evaluations are a critical part of AI progress at Meta Superintelligence Labs, determining what capabilities get built, which features get prioritized, and how quickly our models improve. As a Systems and ML Infrastructure Engineer on this team, you will build the platforms and services that enable reliable evaluation of our most advanced AI models across text, vision, audio, and beyond. You’ll work alongside researchers and engineers to turn rapidly evolving research workflows into scalable, dependable infrastructure.This is a highly technical software engineering role focused on distributed systems, developer infrastructure, and production-grade ML platforms. You will design and own systems for scheduling and executing evaluation workloads, managing datasets and model artifacts, monitoring correctness and performance, and making results reproducible and easy to consume. The infrastructure you build will directly support research decisions and major model lines within MSL, making reliability, scalability, operational excellence, and engineering rigor paramount.You will succeed by moving quickly in an open-ended research environment while building durable systems, reducing operational toil, and creating abstractions that help researchers iterate faster. If you are passionate about building the technical foundation for frontier AI development and thrive in fast-paced, high-impact environments, we encourage you to apply.
Required Skills:
Software Engineer, Systems & ML Infrastructure - MSL FAIR Foundations Responsibilities:
- Design, build, and operate scalable infrastructure for running evaluations across large model fleets, datasets, modalities, and compute environments
- Develop orchestration, scheduling, data, and artifact-management systems that make the evaluation workflows reliable and reproducible
- Build APIs, abstractions, and developer tools that allow researchers to launch, debug, compare, and interpret evaluations efficiently
- Improve system reliability through testing, observability, capacity planning, performance optimization, and automated failure recovery
- Partner with research and engineering teams to translate new evaluation requirements into reusable platform capabilities
Requirements
- 3+ years of software engineering experience in building backend, distributed, data, or machine learning infrastructure
- Proficiency in Python, C++, or another systems programming language
- Experience designing, implementing, and operating reliable services, platforms, or data-processing systems
- Experience independently delivering medium- to large-scale technical projects from design through production operation
- Demonstrated knowledge of software engineering practices, including testing, code review, observability, incident response, and performance analysis
- Ability to work effectively with researchers and engineers and to adapt to rapidly changing requirements, 1. Experience building infrastructure for large-scale machine learning training, inference, evaluation, or data processing
- Experience with distributed compute systems, workflow orchestration, containers, cluster schedulers, or cloud infrastructure
- Experience with performance profiling, resource efficiency, reliability engineering, and production observability
- Familiarity with language model post-training workflows, including supervised fine-tuning, reinforcement learning, evaluation, and inference, and the infrastructure needed to support them at scale
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
What Are Large Language Models?
Why Upskilling And Reskilling is Important For Developers
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production