Software Engineer, Systems ML Engineering

The Meta Game, Inc.
San Francisco, CA, United States
3 days ago
Apply on www.sanfranciscogigs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$183,997.0 - $257,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Systems Engineering C++ (Programming Language) Code Review Nvidia CUDA Computer Programming Computer Engineering Distributed Computing Environment Distributed Systems General-Purpose Computing on Graphics Processing Units Systems Analysis Python (Programming Language)
+15 more
Machine Learning Open Source Technology Tensorflow Azure Machine Learning Software Engineering Software Systems System Programming Strategies of Testing Pytorch Large Language Models Generative AI Information Technology Performance Monitor Machine Learning Operations Data Pipelines

Job description

Meta is seeking a Staff Software Engineer to join the Systems ML Engineering team, focused on building and scaling the infrastructure and software systems that power large-scale machine learning workloads across Meta’s production fleet. In this role, you will architect and own critical components of the ML systems stack, spanning training infrastructure, model serving, distributed computing frameworks, and ML platform tooling. You will work at the intersection of systems engineering and machine learning to drive reliability, performance, and efficiency for some of the world’s most demanding AI workloads, including large language models and generative AI systems., 1. Design and implement scalable ML systems infrastructure components, including distributed training frameworks, model serving pipelines, and ML platform tooling used across Meta’s production AI workloads

  1. Lead technical design and architecture for major initiatives in the ML systems stack, evaluating trade-offs across performance, reliability, and engineering complexity
  2. Identify and resolve performance bottlenecks in distributed ML training and inference systems through instrumentation, profiling, and targeted optimization
  3. Define and drive service level objectives for ML infrastructure services, building dashboards, alerting, and runbooks to reduce mean time to mitigation during incidents
  4. Collaborate with machine learning researchers, product engineers, and infrastructure teams to translate model development requirements into robust, production-grade systems
  5. Leverage AI-assisted development workflows to accelerate implementation, code review, and system analysis, applying sound judgment on when to rely on AI tooling versus deep domain expertise
  6. Mentor other engineers on ML systems best practices, distributed computing patterns, and engineering craft, including AI-native development workflows
  7. Drive adoption of engineering standards across the team, including testing strategies, staged rollout practices using feature flagging and experimentation frameworks, and proactive monitoring
  8. Contribute to roadmap definition and stakeholder alignment for multi-quarter ML infrastructure investments, communicating technical options and trade-offs to both engineering and cross-functional audiences
  9. Conduct thorough code reviews and establish coding standards that improve maintainability and scalability of the ML systems codebase

Requirements

  1. Bachelor’s degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  2. 8+ years of experience in software engineering with a focus on systems software, distributed computing, or ML infrastructure
  3. Experience designing and implementing large-scale distributed systems, including components such as training orchestration, model serving, or data pipeline infrastructure
  4. Experience with performance analysis and optimization of compute-intensive or distributed workloads, including profiling, benchmarking, and bottleneck identification
  5. Experience leading end-to-end delivery of complex technical projects, including cross-team coordination, milestone planning, and risk mitigation
  6. Experience with C++, Python, or equivalent systems programming languages applied to production ML or infrastructure systems, 1. Experience contributing to or maintaining open-source ML systems or distributed computing projects
  7. Experience building or operating ML platform services including experiment tracking, model registries, feature stores, or inference serving infrastructure
  8. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  9. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  10. Experience with ML frameworks such as PyTorch, including distributed training paradigms such as data parallelism, model parallelism, or pipeline parallelism
  11. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  12. Experience with GPU computing, CUDA programming, or accelerator-aware systems optimization for large-scale AI workloads

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.sanfranciscogigs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · World Congress 2026 Europe

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · World Congress 2026 Europe

Videos

See all

Related articles

See all