> Markdown version of [/jobs/ext/329077-data-engineer](https://www.wearedevelopers.com/jobs/ext/329077-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Tyne & Wear - **Location:** Boldon Colliery, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Amazon S3, Computational Fluid Dynamics, Cloud Computing, Information Engineering, Middleware, JSON, Python (Programming Language), Search Technologies, Siemens NX, Parquet, Multi-Cloud, Data Lakes, Pyspark, AWS Data Analytics, Ansys - **Published:** June 5, 2026 - **Apply:** https://www.apply4u.co.uk/jobs/x/38287576/ ## About the Role the AI onboarding pattern document.Co-design the API middleware blueprint with the Cloud Infrastructure Architect. Must-have Principal-level hands-on data engineering on AWS - 7+ years Deep production experience with S3, S3 Tables, Glue, Athena, and OpenSearch (including k-NN / vector search) Built and shipped vector embedding workloads Strong metadata modelling and data taxonomy design experience for scientific or engineering domains Comfort working with Parquet, JSON-LD, and large binary scientific data formats (mesh, time-series, spectra) Python proficiency; PySpark / Glue job tuning experience Nice-to-have / differentiatorsPrior simulation / CAE / HPC data lake experience (Ansys, Siemens NX, BETA CAE, OpenFOAM, etc.)Familiarity with surrogate model training data pipelinesExperience with SageMaker Unified Studio or comparable governed data-mesh tooling (in case of required integration)Multi-cloud data engineering (AWS GCP) experiencePublished or contributed to AWS data architecture patterns or blueprints ## Description Role summary:The overall technical lead and architect. Designs the metadata schema, builds the simulation onboarding pipeline, deploys metadata embedding pipeline and OpenSearch k-NN vector store, and authors data export format spec for AI/ML use case. This role is the deepest technical seat on the engagement: Key responsibilitiesRun the Sprint 1 architecture review of the existing UAT codebase (S3 + Glue + S3 Tables + OpenSearch + Athena) and deliver written gap findings.Design the metadata schema, taxonomy, and field catalogue (Light, Brain, Power).Tune data orchestration - Glue jobs, Athena queries, S3 Tables config, scheduling. Lead the deep-dive technical sessions with analysts on visualization requirements Build and validate the simulation data onboarding pipeline against real data - including the 30 GB-per-run acoustic spectra dataset.Configure and validate the OpenSearch k-NN vector store and the Bedrock embedding pipeline.Author the AI/ML data export format specification and ## Related Videos - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Cloud Vendor Lock-In - Is it just a new version of the Database Abstraction Layers?](https://www.wearedevelopers.com/videos/1185-cloud-vendor-lock-in-is-it-just-a-new-version-of-the-database-abstraction-layers) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)