Data Engineer

Tyne & Wear
Boldon Colliery, UK
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon S3 Computational Fluid Dynamics Cloud Computing Information Engineering Middleware JSON Python (Programming Language) Search Technologies Siemens NX
+6 more
Parquet Multi-Cloud Data Lakes Pyspark AWS Data Analytics Ansys

Job description

Role summary:The overall technical lead and architect. Designs the metadata schema, builds the simulation onboarding pipeline, deploys metadata embedding pipeline and OpenSearch k-NN vector store, and authors data export format spec for AI/ML use case. This role is the deepest technical seat on the engagement: Key responsibilitiesRun the Sprint 1 architecture review of the existing UAT codebase (S3 + Glue + S3 Tables + OpenSearch + Athena) and deliver written gap findings.Design the metadata schema, taxonomy, and field catalogue (Light, Brain, Power).Tune data orchestration - Glue jobs, Athena queries, S3 Tables config, scheduling. Lead the deep-dive technical sessions with analysts on visualization requirements Build and validate the simulation data onboarding pipeline against real data - including the 30 GB-per-run acoustic spectra dataset.Configure and validate the OpenSearch k-NN vector store and the Bedrock embedding pipeline.Author the AI/ML data export format specification and

Requirements

the AI onboarding pattern document.Co-design the API middleware blueprint with the Cloud Infrastructure Architect. Must-have Principal-level hands-on data engineering on AWS - 7+ years Deep production experience with S3, S3 Tables, Glue, Athena, and OpenSearch (including k-NN / vector search) Built and shipped vector embedding workloads Strong metadata modelling and data taxonomy design experience for scientific or engineering domains Comfort working with Parquet, JSON-LD, and large binary scientific data formats (mesh, time-series, spectra) Python proficiency; PySpark / Glue job tuning experience Nice-to-have / differentiatorsPrior simulation / CAE / HPC data lake experience (Ansys, Siemens NX, BETA CAE, OpenFOAM, etc.)Familiarity with surrogate model training data pipelinesExperience with SageMaker Unified Studio or comparable governed data-mesh tooling (in case of required integration)Multi-cloud data engineering (AWS GCP) experiencePublished or contributed to AWS data architecture patterns or blueprints

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on apply4u.co.uk

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

55 sec

Generating ASCII art branding for AI agent interfaces

Chris Heilmann +2 · LIVE

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

2:03 min

Introduction to open table formats built on Parquet

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

2:31 min

Simplifying terminal output with new escape sequences

Ambesh Singh +1 · LIVE

Videos

See all

Related articles

See all