Data Engineer

Nigel Frank International
Greater London, UK
22 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Data Deduplication Data Profiling Parsing Performance Tuning Spatial Data Infrastructures Data Logging Scripting Azure Data Factory Apache Spark SC Clearance Pandas Matplotlib
+3 more
Scikit Learn Tools for Reporting Raster Graphics

Job description

A skilled Data Engineer is required to support a major data-transformation workstream within a clinical screening and diagnostics environment. The project focuses on modernising an end-to-end screening service through modern, data-driven digital capabilities. You will join a small specialist team responsible for analysing complex datasets spread across multiple system instances, resolving data-quality challenges, and shaping future data models and structures.

Requirements

  • Experience establishing import/export patterns, including handling data extracts, schema discovery, incremental loading and normalising data across multiple source systems.
  • Strong capability in data-transformation-heavy pipelines covering profiling, cleansing, standardisation, conformance and final data publishing.
  • Advanced SQL expertise including profiling, joins/merges, deduplication, anomaly detection and performance tuning.
  • Practical Python scripting experience for automation, parsing, rules engines and data-quality checks, using libraries such as Pandas/Polars, scikit-learn or matplotlib.
  • Experience with modern data tooling (e.g., Spark, Azure Data Factory) or the ability to deliver equivalent functionality in code-based environments.
  • Proven experience working with geospatial datasets (vector, raster, GeoJSON, shapefiles), including coordinate systems, spatial data handling and geospatial analysis workflows.
  • Ability to interpret geographical context and aggregate/upscale local or regional geospatial insights into coherent national- or region-level datasets.
  • Experience working with publicly available official datasets (e.g., census boundaries, geographic lookups, deprivation indices, population estimates).
  • Capability to design rules for completeness, validity and consistency, and to implement exception handling and reconciliation flows.
  • Ability to build version-controlled pipelines with deterministic transformations, logging, lineage and full traceability of data changes.
  • Comfortable working in a secure environment with least-privilege access principles, secure storage/transfer practices and handling of sensitive personal data.

Soft Skills

  • Strong communication and collaboration skills.
  • A team-focused, cooperative approach.
  • Enthusiasm, engagement and a positive attitude.
  • Proactivity-comfortable working independently without constant direction.
  • Ability to handle ambiguity and adapt to change effectively.

Nice-to-Have Skills

  • Experience working with healthcare or medical-sector datasets, including patient/episode-style records or longitudinal histories.
  • Experience building automated data-profiling dashboards or reporting frameworks., Candidates must be eligible for UK Security Clearance due to the sensitive nature of the data involved.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:56 min

Open-sourcing a complex parsing library for game data

Johan Hutting Johan Hutting · World Congress 2024

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham · LIVE

6:58 min

Analyzing production code coverage data using pandas

Markus Harrer Markus Harrer · World Congress 2021

Videos

See all

Related articles

See all