Data Engineer

Summarythis
Leeds, UK
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Data Analysis Microsoft Azure Cloud Storage Data Architecture Information Engineering Data Governance Data Transformation Logic Synthesis of Circuits Python (Programming Language) Azure Data Lake SQL Databases Azure Data Factory
+13 more
Apache Spark Git SC Clearance Microsoft Fabric Data Lakes Pyspark Information Technology Data Lakehouse Azure Synapse Analytics Software Version Control Data Pipelines Legacy Systems Databricks

Job description

Design, build, and optimise data pipelines using Microsoft Fabric (Data Factory, Dataflows Gen2) and Azure Data Factory to ingest data from legacy systems and third-party sources.Develop and maintain the Bronze, Silver, and Gold layers of the lakehouse architecture using OneLake, Delta Lake, and Apache Spark within Fabric.Implement data transformation logic using PySpark, SQL, and Fabric Notebooks; ensure data quality, lineage, and cataloguing via Microsoft Purview.Collaborate with Technical Architects and Infrastructure Engineers to support CI/CD pipelines, infrastructure-as-code, and platform automationContribute to knowledge transfer workshops, running instructions, and documentation to build internal capability.Support governance compliance including Digital Design Authority reviews, Red Lines Assessments, and security controls.

Requirements

Essential requirementsStrong experience with Microsoft Azure Fabric (Lakehouses, Data Pipelines, Dataflows Gen2, Fabric Notebooks)Strong experience with Azure Data Factory (ADF) for orchestration and data movementStrong proficiency in PySpark, SQL, and Python for large-scale data transformationStrong experience with Delta Lake, Apache Spark, and OneLake architectureGood knowledge of Microsoft Purview for data governance, cataloguing, and lineageGood experience with Azure DevOps, Git-based version control, and CI/CD pipelinesGood understanding of data lakehouse architecture (medallion architecture - Bronze/Silver/Gold)Good knowledge of Azure storage services (ADLS Gen2, Azure Blob Storage Nice to have skillsExperience with Azure Synapse Analytics or migration from Synapse to FabricFamiliarity with Databricks or equivalent distributed processing platformsExperience in UK public sector or government data environmentsUnderstanding of SC clearance requirements and government security classificationsKnowledge of DDAT frameworks and GDS delivery standards QualificationsRelevant degree in Computer Science, Data Engineering, or related discipline (or equivalent experience)Microsoft Certified: Azure Data Engineer Associate (DP-203) - desirableMicrosoft Fabric Analytics Engineer (DP-600) - desirable

About the company

Job summaryThis is a pivotal engineering role at the heart of one of the most significant data transformation programmes in largest public service department,. You will be joining a strategic engagement programme to design, build, and operationalise a cloud-native data lakehouse on Microsoft Azure Fabric.This is with the UK’s largest public service department, serving over 22 million citizens, and this platform will directly underpin data-driven decision-making at national scale.You will take a leading role in the design and delivery of data pipelines, data transformation layers, and lakehouse infrastructure using Microsoft Fabric, Azure Data Factory, and related Azure-native technologies.You will work in agile squads alongside architects, analysts, and DevOps engineers, contributing to private beta builds, public beta expansion, and full platform operationalisation.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on apply4u.co.uk

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:27 min

Managing traffic and tracking costs with Databricks Unity Catalog

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:50 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

Videos

See all

Related articles

See all