Senior Software Engineer, Training Efficiency

Waymo LLC
Mountain View, CA, United States
16 days ago
Apply on jobs.localjobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$238,000.0 - $302,000.0
Working hours
Regular working hours

Tech stack

C++ (Programming Language) Profiling Distributed Systems Data Flow Control Python (Programming Language) Machine Learning Tensorflow Data Processing Information Technology Machine Learning Operations Data Pipelines

Job description

Waymo is an autonomous driving technology company with the mission to be the world’s most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World’s Most Experienced Driver-to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.

The Waymo ML Infrastructure team works with Research and Production teams to develop models in Perception and Planning that are core to our autonomous driving software. We help our partners by offering the best solutions for the entire model development lifecycle. These solutions are developed in close collaboration with teams at Google. They are geared towards both scaling models and solving problems unique to ML for autonomous driving. You will improve the runtime efficiency of input data pipelines for large-scale training workloads. This is a unique opportunity to work on ML systems and improve on our model training processes.

You Will:

  • Design, and improve distributed input data pipelines for large-scale ML training workloads.
  • Collaborate with researchers and ML engineers to resolve bottlenecks in data pipeline performance.
  • Improve runtime goodput of ML training workload, including optimizing input data processing systems, ensuring scalability and reliability across distributed environments.
  • Implement and maintain advanced ML infrastructure tools, including ML Pathways, Grain, JAX, and TensorFlow.
  • Evaluate and integrate modern technologies to enhance the performance and scalability of ML systems.
  • Promote best practices for distributed systems architecture and contribute to technical leadership within the team.

Requirements

  • B.S. in Computer Science, Math, or 5+ years equivalent real-world experience.
  • Proficient in distributed systems design with an understanding of ML data pipeline optimization.
  • Experience with ML frameworks, including TensorFlow and JAX.
  • Hands-on experience libraries like Grain or tf.data service.
  • Solid programming skills in Python and C++.
  • Practical familiarity with profiling tools to uncover performance bottlenecks.

We Prefer:

  • MS in Computer Science, Math
  • Familiarity with distributed dataflow frameworks like ML Pathways

Benefits & conditions

The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process.

Waymo employees are also eligible to participate in Waymo’s discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements. Salary Range $238,000-$302,000 USD

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.localjobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · World Congress 2024

1:19 min

Advancing autonomous driving capabilities with specialized software talent

Katrin Lehmann Katrin Lehmann +1 · Coffee With Developers

2:17 min

Comparing code profiling with surface level monitoring

Jérôme Vieilledent · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all