Forward Deployed Engineer

Sieve Inc.
San Francisco, CA, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Python (Programming Language) Tensorflow Data Processing Pytorch Free and Open-Source Software Data Pipelines

Job description

As a Forward Deployed Engineer at Sieve, you’ll work on highly specific dataset problems for frontier AI labs. We’re looking for someone with a strong bias to action who likes working closely with customers, untangling messy requirements, and shipping fast.

You’ll work closely with customers and internal teams to understand exactly what data is needed, then turn ambiguous requirements into production systems that can find, generate, filter, transform, evaluate, and package high-quality video datasets at scale.

Requirements

  • Comfortable working directly with customers or external teams to translate ambiguous needs into concrete technical systems
  • Strong Python developer with hands-on experience in PyTorch or similar ML frameworks
  • Experience building custom algorithms, model workflows, or large-scale data pipelines
  • Strong intuition for dataset quality, filtering, labeling, evaluation, and edge cases
  • Able to break customer-level goals down into the models, heuristics, infrastructure, and QA steps needed to deliver
  • Writes clean, maintainable code and can move quickly without creating brittle systems
  • Deep passion for video, media technologies, and frontier AI applications
  • Motivated by delivering end-to-end outcomes, not just training models or writing research code
  • Bonus: Experience with large-scale video, audio, or multimodal data processing
  • Bonus: Active contributor to open source projects
  • Bonus: Experience as an early hire at a startup
  • In-person at our SF HQ

About the company

Sieve is a multi-modal lab curating the world’s highest-quality training datasets - spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet traffic, and across modalities, data has become the enabling medium powering creativity, communication, gaming, AR/VR, and robotics. Sieve exists to solve the biggest bottleneck in the growth of these applications: high-quality training data., Sieve is one of the most capital-efficient teams in AI - roughly 30 people serving the world’s leading AI labs across every major data modality. You’ll join early, own problems end-to-end, and watch your work ship directly into the models defining the frontier.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · World Congress 2024

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · World Congress 2026 Europe

Videos

See all

Related articles

See all