Engineering Manager, Platform Systems - AIMS Engineering New

Netflix, Inc.
Reading, MA, United States
4 days ago
Apply on www.gamesjobsdirect.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Distributed Systems Large Language Models Reliability of Systems Machine Learning Operations Automation Anywhere

Job description

This is a role for a senior engineering leader who can operate at the intersection of technical strategy and organizational execution. You will own the engineering roadmap for Platform Systems, build and scale a high performing team, and set the engineering culture standard for AIMS as a whole. This is not a role for someone who manages from a distance. The problems are technically deep, the dependencies are broad, and the org needs a leader who can credibly engage on distributed systems architecture, steer the design of subsystems that both Netflix’s classic and next generation ML systems depend on, and raise the bar on developer productivity across the org at the same time.

What You’ll Do

  • Define and own the long term technical strategy for Platform Systems, covering reliability, scalability, capacity management, cost efficiency, the subsystems that support next generation AI workflows, and AI stack modernization, and translate it into a roadmap the team and org can execute against.
  • Set the engineering excellence standard for AIMS by owning observability, reliability practices, on call hygiene, and engineering norms that raise the bar across the org.
  • Own the subsystems for observability, evaluation, and tooling that Netflix’s next generation ML architecture depends on. Extend the subsystems into everyday engineering workflows, measurably reducing friction in the development lifecycle.
  • Hire, develop, and retain a team of engineers across a range of seniority levels; create an environment where strong technical contributors grow into senior leaders.
  • Partner at the director level across AIMS and infrastructure teams, representing Platform Systems in strategic planning and advocating for investments with compounding impact across the org.
  • Guide the AI stack modernization effort, establishing the migration path and alignment with partner teams across functions, as one part of the broader technical roadmap.
  • Navigate ambiguity at the intersection of platform engineering, ML infrastructure, and organizational change.

Requirements

  • Significant experience leading platform or infrastructure engineering organizations operating distributed systems at scale.
  • Strong technical depth in distributed systems. You can engage credibly on architecture, tradeoffs, and design decisions with your most senior engineers.
  • Hands-on experience building subsystems that support advanced agentic architectures, such as memory, trace, eval, and replay pipelines, or orchestration and routing layers for complex model systems. This should go beyond wiring up an LLM API call.
  • Experience with ML infrastructure at scale. You understand the operational complexity of online and offline model training pipelines and the platform demands they place on data, compute, and serving infrastructure.
  • Track record of building high performing engineering teams and developing senior technical talent.
  • Track record of improving system reliability, reducing infrastructure costs, and accelerating engineering velocity at scale.
  • Developer productivity as a foundational instinct. You default to removing friction, automating toil, and investing in the tools and systems that make engineers faster.
  • Operates effectively at the intersection of engineering strategy and organizational leadership; comfortable influencing without direct authority across a large org.
  • High tolerance for ambiguity and competing priorities in a fast moving environment., * Demonstrated experience driving large, cross functional technical programs, such as infrastructure modernization or platform migrations, from strategy through execution, without disrupting product delivery.
  • Familiarity with modern AI/ML infrastructure patterns including feature stores, model serving platforms, and experiment frameworks.
  • Familiarity with personalization systems or recommendation platforms.

Benefits & conditions

Generally, our compensation structure consists solely of an annual salary; we do not have bonuses. You choose each year how much of your compensation you want in salary versus stock options. To determine your personal top of market compensation, we rely on market indicators and consider your specific job family, background, skills, and experience to determine your compensation in the market range. The range for this role is $523,000.00 - $920,000.00. This compensation range will vary based on location.

Netflix provides comprehensive benefits including Health Plans, Mental Health support, a 401(k) Retirement Plan with employer match, Stock Option Program, Disability Programs, Health Savings and Flexible Spending Accounts, Family-forming benefits, and Life and Serious Injury Benefits. We also offer paid leave of absence programs. Full-time hourly employees accrue 35 days annually for paid time off to be used for vacation, holidays, and sick paid time off. Full-time salaried employees are immediately entitled to flexible time off. See more details about our Benefits here.

Netflix is a unique culture and environment. Learn more here.

About the company

At Netflix, our mission is to entertain the world. Together, we are writing the next episode - pushing the boundaries of storytelling, global fandom and making the unimaginable a reality. We are a dream team obsessed with the uncomfortable excitement of discovering what happens when you merge creativity, intuition and cutting-edge technology. Come be a part of what’s next.

About the Team

AI for Member Systems (AIMS) runs the AI systems behind every recommendation, search result, and personalized experience for 300M+ members. The stack powering it is large and battle tested, built to meet the demands of its time, and remarkably effective at doing so. But AI/ML is moving fast, and the infrastructure that got us here needs to evolve to meet what’s next: new model paradigms, tighter cost and efficiency expectations, and the operational maturity that comes with running AI at this scale.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.gamesjobsdirect.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:32 min

Fundamentals and limitations of large language models

Krzystof Czieslak · LIVE

1:29 min

Overcoming challenges in AI-assisted distributed system development

Przemysław Ładyński Przemysław Ładyński · World Congress 2026 Europe

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

2:25 min

Understanding the evolution and nature of large language models

Krzysztof Cieślak Krzysztof Cieślak · World Congress 2024

6:26 min

Bringing accurate time synchronization to global distributed systems

Werner Vogels Werner Vogels · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all