Software Engineer, Systems ML (Technical Leadership)

The Meta Game, Inc.
Sunnyvale, CA, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Abstraction Layers Artificial Intelligence C++ (Programming Language) Nvidia CUDA Python (Programming Language) Machine Learning Software Engineering Information Technology Machine Learning Operations

Job description

Experteer Overview As a Principal Software Engineer in Systems ML Engineering, you shape the architectural foundations of large-scale ML infrastructure. You define multi-year roadmaps and drive cross-team execution to deliver training, inference, and compiler optimization capabilities. You’ll tackle the toughest cross-system ML infrastructure challenges and enable AI-native workflows that amplify engineering impact. You work closely with research, hardware, and product teams to translate advances into production, with a strong emphasis on reliability and performance. This role offers an opportunity to influence technical strategy at scale and help Meta stay at the forefront of AI-driven infrastructure. Compensation / Benefits * Identify and solve complex cross-system ML infrastructure challenges across training, inference, compiler optimization, and hardware-software co-design * Define extensible architectural standards and foundations ensuring consistency and reliability across multiple orgs * Own multi-year technical roadmap for ML systems infrastructure balancing short-term delivery and long-term platform health * Leverage AI-native tooling to reduce engineering toil and accelerate cross-disciplinary work * Drive performance improvements for large-scale ML training/inference systems across subsystems and abstraction layers * Establish invariants, correctness proofs, and systemic reliability practices to prevent failures * Collaborate with research, hardware, and product teams to translate ML advances into production gains * Assess emerging AI and computing technologies and influence organizational strategy * Mentor engineers, lead programs, and foster a culture of rigor and craftsmanship in ML systems Tasks * Bachelor in Computer Science (or related field) and 12+ years in software engineering with ML systems specialization * Experience architecting and delivering large-scale ML training or inference infrastructure with measurable cross-team impact * Proven track record leading multi-year cross-functional initiatives with metrics, dependencies, and cross-org execution * Proficient in high-performance ML systems infrastructure using C++, Python, or CUDA * Experience influencing technical direction across multiple teams via proposals, design reviews, and stakeholder alignment * Familiarity with ML compiler stacks or hardware-software co-design for ML accelerators Key requirements *

Requirements

_ orgs * Own multi-year technical roadmap for ML systems infrastructure balancing short-term delivery and long-term platform health * Leverage AI-native tooling to reduce engineering toil and accelerate cross-disciplinary work * Drive performance improvements for large-scale ML training/inference systems across subsystems and abstraction layers * Establish invariants, correctness proofs, and systemic reliability practices to prevent failures * Collaborate with research, hardware, and product teams to translate ML advances into production gains * Assess emerging AI and computing technologies and influence organizational strategy * Mentor engineers, lead programs, and foster a culture of rigor and craftsmanship in ML systems Tasks * Bachelor in Computer Science (or related field) and 12+ years in software engineering with ML systems specialization * Experience architecting and delivering large-scale ML training or inference infrastructure with measurable cross-team impact * Proven track aaa and leading multi-year cross-functional initiatives with metrics, dependencies, and cross-org execution * Proficient in high-performance ML systems infrastructure using C++, Python, or CUDA * Experience influencing technical direction across multiple teams via proposals, design reviews, and stakeholder alignment * Familiarity with ML compiler stacks or hardware-software co-design for ML accelerators Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

1:23 min

Handling compatibility and abstraction layers in composable systems

Loïc Carbonne Loïc Carbonne · WWC 2024

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

Videos

See all

Related articles

See all