> Markdown version of [/jobs/ext/2720453-data-engineer](https://www.wearedevelopers.com/jobs/ext/2720453-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Toyota Research Institute - **Location:** Los Altos, United States - **Experience:** Expert - **Salary:** $258,750.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Amazon S3, C++ (Programming Language), Cloud Engineering, Code Review, Continuous Integration, Data Discovery, Information Engineering, Data Infrastructure, Extract Transform Load (ETL), Serialization, Software Debugging, Global Positioning Systems (GPS), Protocol Buffers, Python (Programming Language), Machine Learning, Software Deployment, Software Systems, Spatial Data Infrastructures, SQL Databases, Unstructured Data, Management of Software Versions, Web Application Frameworks, Data Logging, Google Cloud, Cloud Platform System, Data Ingestion, Apache Spark, Indexer, Build Management, Kubernetes, Information Technology, Lidar, Data Pipelines, Data Generation - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-data-engineer-toyotaresearch-7073319 ## About the Role * Bachelor's or Master's degree in Computer Science, Data Engineering, or a related field. * 8+ years of experience building data-intensive software systems, ideally in robotics, autonomous driving, or large-scale ML environments. * Proficient in Python, SQL, and familiar with C++. * Experience designing ETL pipelines using modern frameworks (e.g., Apache Spark, Flyte, Union). * Strong knowledge of cloud-native architectures, including AWS services (e.g., S3, or equivalents (Google Cloud platform) * Familiarity with sensor data types (camera, lidar, radar, GPS/IMU) and common data serialization formats (e.g., protobuf. ROS2bag, MCAP). * Deep understanding of data quality, observability, and lineage in high-volume systems. * Track record of building reliable and performant infrastructure that supports both ad-hoc exploration and repeatable production workflows., * Experience in AD/ADAS, robotics, or autonomous systems - especially handling perception or planning datasets. * Familiarity with ML pipeline orchestration frameworks (e.g. Kubeflow, SageMaker, etc). * Experience working with temporal or spatial data, including geospatial indexing and time-series alignment. * Exposure to synthetic data generation, simulation logging, or scenario replay pipelines. * Strong software engineering fundamentals, CI/CD, testing, code review, and service deployment best practices. ## Description * Design and implement scalable, production-grade pipelines for data ingestion, transformation, storage, and retrieval from vehicle fleets and simulation environments. * Build internal tools and services for data labeling, curation, indexing, and cataloging across large and diverse datasets. * Collaborate with ML researchers, autonomy engineers, and data scientists to design schemas and APIs that power model training, evaluation, and debugging. * Develop and maintain feature stores, metadata systems, and versioning infrastructure for structured and unstructured data. * Support the generation and integration of synthetic datasets with real-world logs to enable hybrid training and simulation workflows. * Optimize pipelines for cost, latency, and traceability, ensuring reproducibility and consistency across environments. * Partner with simulation and cloud platform teams to automate workflows for closed-loop testing, scenario mining, and performance analytics. ## Related Videos - [How to develop an autonomous car end-to-end: Robotic Drive and the mobility revolution](https://www.wearedevelopers.com/videos/22-how-to-develop-an-autonomous-car-end-to-end-robotic-drive-and-the-mobility-revolution) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Dynamic Entities in .NET: Building Low-Code Systems on Top of Entity Framework Core](https://www.wearedevelopers.com/videos/100218-dynamic-entities-in-net-building-low-code-systems-on-top-of-entity-framework-core) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)