> Markdown version of [/jobs/ext/3251226-embedded-data-engineer](https://www.wearedevelopers.com/jobs/ext/3251226-embedded-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Embedded Data Engineer - **Company:** The Trainline - **Location:** UK - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Data Analysis, Information Engineering, Data Infrastructure, Data Mart, Data Transformation, Data Systems, Distributed Computing Environment, Github, Python (Programming Language), Machine Learning, Operational Databases, SQL Databases, Parquet, Feature Engineering, Apache Spark, Build Management, Containerization, Terraform, Data Pipelines, Docker, Jenkins - **Published:** September 9, 2026 - **Apply:** https://startup.jobs/embedded-data-engineer-ml-trainlinegroup-com-9980447 ## About the Role * Working knowledge of Python and SQL. * Experience building data pipelines for downstream machine learning workloads, including feature engineering and model training workflows. * Comfort with data modelling and building efficient data marts and warehouses in the cloud. * Experience building data pipelines using tools such as Spark and Airflow, or similar technologies, within a cloud environment such as AWS. * Familiarity with both real-time and batch data workloads, along with modern data transformation and orchestration patterns. * Ideally, you may also have experience with parallel or distributed training frameworks such as Ray, or with modern data formats such as Parquet and Iceberg. * It would also be helpful if you have some experience with Infrastructure as Code (Terraform) and containerisation (Docker) to support automated, standardised deployments. * You may also have contributed to or maintained CI/CD pipelines (such as Jenkins or GitHub Actions) as part of production grade data systems, and enjoy solving complex data problems collaboratively. ## Description As an Embedded Data Engineer in ML, you'll sit within the Machine Learning team, working day to day with Machine Learning Engineers and Data Scientists to build reliable datasets for ML use cases. You'll also have access to Trainline's wider Data Engineering, Data Platform, and analytics community, working alongside other embedded Data Engineers in ML, including senior and principal engineers. In this role as the Embedded Data Engineer (ML), you will... * Design and build scalable data pipelines, data models, and feature stores that support analytics and machine learning workloads within the ML domain. * Deploy and maintain cloud-native data applications on AWS, using CI/CD pipelines to automate builds, testing, and releases. * Maintain the technical quality, performance, and reliability of production data pipelines through strong observability and engineering best practices. * Collaborate closely with Machine Learning Engineers and Data Scientists to build reliable, well-structured datasets that power ML use cases. * Work with the wider Data Engineering, Data Platform, and analytics community to share knowledge and align on best practices across teams.