Software Engineer, Data Systems

SF MISSION, LLC
San Francisco, CA, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon S3 Apache HTTP Server C++ (Programming Language) Cloud Computing Nvidia CUDA Databases Information Engineering Data Systems Distributed Data Store Distributed Systems Search Technologies
+9 more
Database Engines Data Streaming Parquet Graphics Processing Unit (GPU) Data Storage Technologies Large Language Models Indexer Data Lakes Data Pipelines

Job description

Our goal is to build Scenario Mining and Data Curation for robot fleet data. We empower Physical AI and robotics teams to instantly find, curate, and stream the data they need to train frontier models.

Eventual is an agile team where every engineer has high ownership across the stack from our compute infrastructure, to our data storage/querying layers and model training/deployment., As a Software Engineer on the Data Systems team, you will build key capabilities for Eventual. We build storage and a high-throughput data engine over petabytes of video, lidar and high-frequency telemetry data to power frontier robotics and Physical AI labs. You will work directly on the architecture powering real-time indexing of perception/robotics data, distributed storage/compute, and dataloading at line-rate to GPUs for model training and inference. We operate as a tight-knit, experienced engineering team that values technical autonomy, deep execution, and a passion for solving hard distributed systems problems., * Multimodal Storage (Data Lake): built against modern columnar data lake formats (Apache Parquet, Apache Iceberg etc) optimized for high-dimensional video, lidar and sensor logs.

  • Query Engine: build powerful querying capabilities. Our multi-stage query systems are built on database fundamentals such as partitioning, indexing, query planning, embeddings/vector search for search and retrieval as well as LLMs/VLMs for perception-based query predicates.
  • Dataloading: Improve memory stability, throughput, and zero-copy data flow through streaming computation and line-rate CUDA tensor delivery to GPUs.

Requirements

  • Proven track record building resilient, high-throughput distributed systems or database engines using Rust or C++.
  • 3+ years of experience diving deep into engine internals-such as vectorized execution, query planning/optimization, distributed task scheduling, or zero-copy networking.
  • Practical exposure to scaling cloud infrastructure (AWS S3) and managing heavy-compute data pipelines (bonus points for experience with CUDA, GPU streaming, or video decoding frameworks).
  • High agency and adaptability to thrive in an autonomous, fast-paced startup environment building cutting-edge infrastructure for frontier robotics

About the company

From humanoid robots to autonomous vehicles, every Physical AI model is trained on petabytes of video, lidar, radar, and sensor data. Today’s data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not video corpora. And understanding that video still means paying a person to watch it, ten dollars an hour of footage at the low end. So teams check a sample and hope it represents the rest. The footage grows every year; the budget to look at it doesn’t.

Eventual was founded in 2022 to close that gap. Our open-source engine, Daft, is purpose-built for multimodal AI: 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at companies like Mobileye, TogetherAI. On top of it we’re building the infrastructure that finds any situation you can describe across a fleet’s entire video history, and turns it into a training set or an alert someone can still act on. We fine-tune and run the vision models ourselves, which makes indexing every hour cheaper than annotating a sample.

We’re building this with the top Physical AI labs and GPU cloud providers. We’ve raised $30M from investors like Felicis, CRV, Y Combinator, and angels from the co-founders of Databricks and Perplexity. Our team comes from AWS, Lyft, and Tesla. We powered the last generation of Physical AI in self-driving; now we’re doing it for the next.

Join our small (but powerful!) team, 4 days/week in our SF Mission District office.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:36 min

Analyzing limitations with PostgreSQL bitmap heap scans

Dharin Shah Dharin Shah · World Congress 2025

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:32 min

Generating functional runtime database columns using indexer properties

Halil İbrahim Kalkan Halil İbrahim Kalkan · World Congress 2026 Europe

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all