> Markdown version of [/jobs/ext/2986080-data-ingestion-engineer](https://www.wearedevelopers.com/jobs/ext/2986080-data-ingestion-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Ingestion Engineer - **Company:** WAYVE LLC - **Location:** Sunnyvale, CA, United States (Remote available) - **Salary:** $210,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Airflow, Batch Processing, Big Data, Data Infrastructure, Extract Transform Load (ETL), Software Debugging, Distributed Data Store, Distributed Systems, Python (Programming Language), Operational Databases, Queue Management Systems, Data Processing, Data Ingestion, Apache Spark, Data Lakes, Data Management, Data Pipelines, Databricks - **Published:** September 18, 2026 - **Apply:** https://www.thejobnetwork.com/job/477f6fe2-b816-4411-a389-ffaf4f94cb88/swe-data-ingestion ## About the Role You are an experienced Data Engineer, Platform Engineer or Distributed Systems Engineer who enjoys working on large-scale production data systems. You have seen how data pipelines behave in the real world: messy inputs, strange edge cases, corrupt files, stalled queues, unexpected formats and failures that only appear at scale. You are comfortable digging into those problems, finding the root cause and making systems better as a result. You combine strong technical depth with a practical, collaborative approach. You can take ownership of complex systems, work effectively across teams and balance urgent operational needs with thoughtful, durable engineering improvements. **Essential** - Strong production experience with Apache Spark. - Strong Python engineering experience. - Experience building, debugging or operating large-scale data-ingestion, ETL or data-processing pipelines. - Experience with distributed data-processing systems. - Ability to optimise jobs for throughput, compute efficiency and reliability. - Experience debugging production pipeline failures. - Comfort working with messy, corrupt, incomplete or inconsistent data. - Understanding of orchestration across multi-step pipelines and downstream dependencies. - Ability to work independently in a fast-moving, highly technical environment. - A practical, delivery-focused mindset with a focus on continuous improvement. - Experience working at significant data scale, ideally PB-scale or similarly high-throughput environments. **Desirable** Experience in one or more of the following areas would be a strong advantage: - Airflow, Flyte, Databricks Workflows or similar orchestration tooling. - Databricks, Delta Lake or Delta tables. - Scala or Java, especially in Spark-based environments. - Queue-based processing, retry handling and priority data workflows. - High-throughput batch data-processing systems. - Production systems with many data producers, consumers or external data sources. - Handling third-party, partner or supplier data with inconsistent formats and quality issues. - Automotive, robotics, autonomy, mapping, ML data platforms or embodied AI environments. - Cost optimisation for compute- and storage-heavy data platforms. - High-performance engineering experience from domains such as trading, where it includes relevant distributed-systems or throughput-focused work. ## Description You will work within the Data Ingestion team to improve the reliability, efficiency and throughput of the pipelines that move real-world driving data through Wayve. - Debug and resolve failing or blocked ingestion pipelines. - Investigate issues caused by corrupt, malformed or unexpected data. - Design and implement more resilient pipelines so individual bad data segments do not block wider workflows. - Improve how we handle varied data formats from partners, suppliers and third-party sources. - Support orchestration across multi-step ingestion workflows, including dependencies, retries and queue management. - Optimise Spark jobs and data-processing pipelines for throughput, compute efficiency and reliability. - Reduce operational toil around failed jobs, stalled pipelines and manual interventions. - Work on high-volume batch-processing systems where throughput, reliability and cost all matter. - Help prioritise and unblock important datasets for downstream annotation, data science and model training teams. - Partner with engineers across Data Platform and downstream teams to deliver both immediate improvements and scalable long-term solutions. - Contribute to the technical direction, maintainability and operational excellence of Wayve's ingestion platform. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)