Senior Machine Learning Engineer, AI Platform

Mozilla Corporation
San Francisco, CA, United States
3 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Compensation
$163,000.0 - $218,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Systems Engineering Profiling Code Review Software Debugging Distributed Computing Environment Distributed Systems Python (Programming Language) Machine Learning Open Source Technology Performance Tuning Azure Machine Learning
+11 more
Strategies of Testing Management of Software Versions Data Logging Cloud Platform System Backend AI Platforms Kubernetes Low Latency Deployment Automation Machine Learning Operations Docker

Job description

The AI Platform team is responsible for building the foundational infrastructure that powers intelligent experiences across Mozilla products. This includes model training pipelines, high-throughput inference services, GPU orchestration, and secure, privacy-respecting AI systems that operate reliably at global scale., We’re looking for a Machine Learning Engineer with a strong platform mindset to help design, build, and operate Mozilla’s AI platform. In this role, you’ll work at the intersection of machine learning, distributed systems, and production infrastructure-ensuring that models can be trained, deployed, and served efficiently, securely, and at scale. You will collaborate closely with product, infrastructure, and security teams to enable fast iteration while meeting strict performance and privacy requirements.

What You’ll Do:

  • Design, build, and operate core AI platform components used to train, deploy, and serve machine learning models in production environments.
  • Own model serving and inference workflows end-to-end, driving improvements in reliability, scalability, performance, and operational excellence.
  • Lead efforts to optimize inference systems for throughput, latency, and cost efficiency across CPU and GPU workloads.
  • Design and manage GPU-based inference and training workloads, including performance tuning, capacity planning, and resource utilization optimization.
  • Own and improve critical parts of the model lifecycle, including packaging, versioning, testing strategies, validation, and deployment automation.
  • Implement and evolve observability practices (metrics, logging, tracing, alerting) to improve visibility and operational resilience of ML services and pipelines.
  • Partner closely with product, infrastructure, security, and data teams to design scalable platform capabilities that enable AI-powered features.
  • Contribute to technical design discussions, propose architectural improvements, and mentor junior engineers through code reviews and knowledge sharing.
  • Participate in and help improve operational processes, including incident response, on-call rotations, and post-incident reviews.

Requirements

Do you have experience in Systems engineering?, Do you have a Master’s degree?, * Bachelor’s degree with 4-6 years of relevant industry experience, or Master’s degree with significant hands-on experience building and operating production ML systems, or work experience equivalent

  • Strong experience developing in Python for machine learning systems, backend services, or distributed data processing.
  • Proven experience deploying and operating ML workloads in cloud environments, including production-grade infrastructure.
  • Solid understanding of model serving architectures, inference pipelines, and performance tradeoffs (latency, throughput, cost, scaling strategies).
  • Hands-on experience working with GPU-based workloads and accelerated computing in production settings.
  • Experience designing CI/CD pipelines and development workflows that support reliable ML system deployment.
  • Ability to independently scope and drive technical initiatives while balancing product and operational priorities.
  • Strong problem-solving skills and the ability to debug performance and reliability issues in distributed systems.
  • Clear and effective communication skills, with experience collaborating across engineering, product, and infrastructure teams.

Bonus Skills:

  • Experience implementing inference optimization strategies such as batching, quantization, compilation, model conversion, or hardware-specific tuning.
  • Familiarity with containerization and orchestration systems (e.g., Docker, Kubernetes) in production environments.
  • Experience designing observability systems for distributed services, including metrics strategy and performance profiling.
  • Exposure to privacy-preserving ML techniques, security best practices, or responsible AI system design.
  • Contributions to open-source ML infrastructure projects or leadership in building reusable internal ML tooling.

Benefits & conditions

4.14.1 out of 5 stars San Francisco, CA Remote $163,000 - $218,000 a year, Pulled from the full job description

  • Work from home stipend
  • Referral program
  • Paid parental leave
  • AD&D insurance
  • Parental leave
  • Health insurance
  • Vision insurance, * Generous performance-based bonus plans to all eligible employees - we share in our success as one team
  • Rich medical, dental, and vision coverage
  • Generous retirement contributions with 100% immediate vesting (regardless of whether you contribute)
  • Quarterly all-company wellness days where everyone takes a pause together
  • Country specific holidays plus a day off for your birthday
  • One-time home office stipend
  • Annual professional development budget
  • Quarterly well-being stipend
  • Considerable paid parental leave
  • Employee referral bonus program
  • Other benefits (life/AD&D, disability, EAP, etc. - varies by country)

About Mozilla

Mozilla exists to build the Internet as a public resource accessible to all because we believe that open and free is better than closed and controlled. When you work at Mozilla, you give yourself a chance to make a difference in the lives of Web users everywhere. And you give us a chance to make a difference in your life every single day. Join us to work on the Web as the platform and help create more opportunity and innovation for everyone online.

About the company

Mozilla Corporation is the non-profit-backed technology company that has shaped the internet for the better over the last 25 years. We make pioneering brands like Firefox, the privacy-minded web browser, and Pocket, a service for keeping up with the best content online. Now, with more than 225 million people around the world using our products each month, we’re shaping the next 25 years of technology and helping to reclaim an internet built for people, not companies. Our work focuses on diverse areas including AI, social media, security and more. And we’re doing this while never losing our focus on our core mission - to make the internet better for people.

The Mozilla Corporation is wholly owned by the non-profit 501(c) Mozilla Foundation. This means we aren’t beholden to any shareholders - only to our mission. Along with thousands of volunteer contributors and collaborators all over the world, Mozillians design, build and distribute open-source software that enables people to enjoy the internet on their terms.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

53 sec

Creating an open ecosystem for artificial intelligence models

Chris Heilmann +2 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all