> Markdown version of [/jobs/ext/2997358-data-engineer-data-quality-provenance](https://www.wearedevelopers.com/jobs/ext/2997358-data-engineer-data-quality-provenance). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer, Data Quality & Provenance - **Company:** Wayve - **Location:** Leonberg, Germany (Remote available) - **Contract:** Permanent contract - **Skills:** Geographic Information Systems, Application Programming Interfaces (APIs), Airflow, Data Infrastructure, Distributed Computing Environment, Distributed Data Store, Python (Programming Language), Machine Learning, Operational Databases, Standard Sql, Management of Software Versions, Workflow Management Systems, Cloud Platform System, Data Ingestion, Apache Spark - **Published:** September 19, 2026 - **Apply:** https://www.adzuna.de/details/5889560096 ## About the Role * Strong hands-on Python and SQL skills, with solid production software-engineering fundamentals. * Experience designing and operating large-scale distributed data systems, beyond small-scale analytics or reporting pipelines. * Hands-on experience with distributed processing and workflow orchestration technologies, such as Spark, Flyte, Airflow or equivalent tools. * Experience building and operating cloud-based data platforms using object storage, including data organisation, versioning, querying, governance and cost management. * Proven ownership of data quality, lineage, observability, reproducibility and incident response for production data workflows. * Experience translating ambiguous requirements from ML, data-science, robotics or similarly technical teams into durable, reusable platform capabilities. * Comfort operating in ambiguity and helping define the boundaries, standards and ways of working for a growing data platform. Desirable * Experience in autonomous vehicles, ADAS, robotics, mapping, drones or another sensor-rich domain. * Familiarity with time-synchronised sensor data, geospatial data, or multimodal datasets. * Understanding of ML training, evaluation, simulation or closed-loop development workflows. ## Description As a Data Engineer focused on Data Quality & Provenance, you will build the data foundation that enables Wayve's autonomous-driving development. You'll turn vast volumes of fleet and simulation data into trusted, discoverable and reproducible datasets that ML, autonomy, simulation and safety teams can use with confidence. This is a high-impact opportunity to define the data products, standards and operating model behind embodied intelligence at petabyte scale., * Design, build and operate scalable batch and streaming pipelines for multimodal fleet and simulation data. * Create data models, catalogs, indexes and query capabilities that make sensor, vehicle-state, map and event data easy to discover and use. * Build versioned, reproducible datasets for training, evaluation, replay, scenario mining and safety analysis. * Develop robust workflows for data ingestion, synchronisation, transformation, curation, labelling and data-quality validation. * Partner with autonomy, ML, simulation and safety engineers to define schemas, APIs and data contracts. * Establish strong standards for lineage, observability, access controls, retention and cost management across the data platform. * Improve the performance, reliability and unit economics of large-scale storage and compute workloads. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Let's Get Aggregated: Custom UDAFs in Spark ](https://www.wearedevelopers.com/videos/1649-let-s-get-aggregated-custom-udafs-in-spark) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)