ML Infrastructure Engineer

Zipline
South San Francisco, CA, United States
11 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$160,000.0 - $250,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Cloud Computing Data Files Data Infrastructure Programming Tools Distributed Computing Environment Python (Programming Language) Machine Learning Cloud Services Tensorflow Software Engineering
+12 more
Software Systems Data Ingestion Pytorch Deep Learning Cloudformation Kubernetes Infrastructure Automation Frameworks Performance Monitor Data Management Machine Learning Operations Terraform Data Pipelines

Job description

As an ML Training & Inference Infrastructure Engineer on the Data Platform team you will be building and scaling the systems powering our data flywheel. This person will work at the intersection of autonomy and the infrastructure, owning systems that make ML development faster, reproducible, observable, and safe.

This role is for a strong software engineer who enjoys the full ML development cycle: data ingestion, processing pipelines, dataset management, distributed training, continuous model integration, evaluation, and deployment. The ideal candidate has strong production engineering habits and is excited to build infrastructure that helps real autonomous systems improve over time.

What You’ll Do

  • Build and operate software infrastructure that enables learning algorithms to leverage Zipline’s large-scale (quickly growing!) fleet data.
  • Design scalable, maintainable data and ML infrastructure for autonomy teams, including dataset creation, validation, training, evaluation, and deployment.
  • Own and improve data pipelines that feed into the ML development loop.
  • Identify and mitigate bottlenecks in the ML development cycle, especially around orchestration, performance, and reproducibility to increase the rate at which we can improve and scale the delivery experience.

Requirements

  • 3+ years of professional software engineering experience, ideally including ML infrastructure, data infrastructure, robotics, autonomy, aerospace, medical devices, or another safety-critical hardware/product environment.
  • Strong software engineering practices in Python in a production setting; comfort designing APIs, services, schemas, jobs, and operational workflows.
  • Experience building reproducible data pipelines and machine-learning pipelines.
  • Experience monitoring data statistics, system performance metrics, pipeline failures, and model/evaluation signals.
  • Working knowledge of ML concepts such as datasets, training, evaluation, optimization, statistics, and modern deep learning workflows.
  • Generalist mindset and willingness to work across cloud services, data platforms, developer tooling, and embedded/robotics-adjacent constraints.
  • Experience with PyTorch or similar ML frameworks.
  • Strong ownership, clear communication, and interest in building secure systems for mission-critical workflows.
  • Experience with Kubernetes or other container orchestration systems for production workloads.
  • Experience with cloud and on-premise production infrastructure, preferably AWS, and infrastructure-as-code tools such as Terraform or CloudFormation., * Experience deploying or evaluating ML systems on real robots, autonomous vehicles, drones, or other hardware products.
  • Experience with large-scale training systems, feature stores, data/versioned artifact stores, model registries, or experiment tracking.
  • Experience with annotation systems, dataset inspection tooling, or active-learning workflows.

Benefits & conditions

The starting cash ranges for this role is $160,000 - $250,000. Please note that this is a target, starting cash range for a candidate who meets the minimum qualifications for this role. The final cash pay for this role will depend on a variety of factors, including a specific candidate’s experience, qualifications, skills, working location, and projected impact. The total compensation package for this role may also include: equity compensation; discretionary annual or performance bonuses; sales incentives; benefits such as medical, dental and vision insurance; paid time off; and more. This position may be hired at one of three levels based on the candidate’s experience, skills, and demonstrated scope:

  • ML Infrastructure Engineer II: $160,000 - $190,000
  • Senior ML Infrastructure Engineer: $190,000 - $220,000
  • Staff ML Infrastructure Engineer: $220,000 - $250,000

Zipline is an equal opportunity employer and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws or our own sensibilities.

We value diversity at Zipline and welcome applications from those who are traditionally underrepresented in tech. If you like the sound of this position but are not sure if you are the perfect fit, please apply!

About the company

Zipline is the world’s largest and most experienced drone delivery service. We are on a mission to serve all humans equally by ensuring access to food, medicine and essential goods anytime, anywhere. We design, build, and operate the world’s largest autonomous logistics system, delivering critical supplies quickly and reliably. Today, Zipline operates on four continents, makes a delivery somewhere in the world every 30 seconds, and has completed millions of deliveries to date, including blood, vaccines, medical supplies, food, and retail products.

Our customers include the world’s largest and most prominent healthcare systems, governments, retailers, restaurants and global businesses who rely on us to save lives, reduce emissions, increase economic opportunity, and provide delivery from point A to point B as fast as possible. The drone is only 15% of what we’ve built to enable seamless, reliable, global operations.

Our system strengthens supply chains, reduces congestion, and gives people time back. With more than 140 million commercial autonomous miles safely flown, Zipline is redefining access to healthcare, consumer products, and food across the globe.

We operate at a global scale and are looking for practical problem solvers who thrive on real-world challenges and rapid growth. Our team is motivated by building systems that have a direct, meaningful impact on people’s lives and by scaling the future of logistics. We are seeking people who sculpt from first principles, enjoy facing adversity, and can do the impossible at record breaking speeds.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:23 min

Exploring specialized career paths within the data science ecosystem

Julian Joseph · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

3:38 min

Reusing email software standards for HTTP file uploads

Imran Nazar · World Congress 2023

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

2:32 min

Overview of Terraform and Terraform Cloud features

Devlin Duldulao · LIVE

Videos

See all

Related articles

See all