Lead Data Architect

Collage Recruitment
Greater London, UK
14 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Artificial Intelligence Automation of Tests Spreadsheets Continuous Integration Data Architecture Python (Programming Language) Microsoft SQL Server Standard Sql SQL Stored Procedures SQL Databases Cloud Platform System Apache Spark
+3 more
Pyspark Infrastructure Automation Frameworks Databricks

Job description

Our client, a leading provider of benchmark data and analytics for the financial services sector, is hiring the first senior engineering leader for their UK Insurance data business. This is a newly created role, not a backfill.

This is a rare opportunity to own a platform transformation from the ground up: migrating a legacy SQL benchmarking engine onto a modern Databricks Lakehouse, and then building the product tooling on top of it.

Our client provides benchmark data behind pricing, acquisition, renewal, and retention decisions across the UK insurance market, built on a proprietary, behavioural dataset that underpins trusted relationships with senior commercial decision-makers across the industry.

You’ll report directly to the Head of Insurance, based in London, and will be the senior technical voice for the UK Insurance business - with the option to line-manage an existing mid-level engineer as the team grows. You’ll also work closely with a broader, global engineering function on platform standards, while owning delivery locally., * Lead the migration of legacy SQL solutions into Databricks Jobs, Workflows, and PySpark/SQL pipelines, starting with a lift-and-shift of core benchmark data, while maintaining operational SLAs throughout

  • Define the target-state data architecture (Lakehouse, Medallion, domain-oriented data products) to support reporting, analytics, and future AI/ML use cases
  • Work directly with business leadership and analysts to codify canonical metric logic - replacing spreadsheet calculations with documented, testable code
  • Own UK Unity Catalog governance: catalogue, schema, table design, ownership, and access control
  • Build automated QA and reconciliation between legacy and new outputs, with the observability to give internal and external stakeholders confidence in the data
  • Implement CI/CD and Infrastructure as Code for Databricks assets
  • Tune Spark and SQL workloads for performance and cost
  • Act as the primary link between UK delivery and the centralised platform engineering function
  • Provide technical leadership and mentorship, with the option to line-manage a mid-level engineer

Requirements

  • Hands-on experience delivering a Databricks migration at production scale - this is the core of the role
  • Strong stakeholder and relationship management skills. You’ll work closely with business leadership, analysts, and engineers across regions, often with ambiguity to navigate
  • Advanced Python, PySpark, and SQL
  • A track record migrating SQL Server workloads, stored procedures, and scheduled jobs to modern, cloud-native platforms
  • Solid understanding of Unity Catalog governance, security models, and access control
  • Experience with CI/CD, automated testing, and Infrastructure as Code, * Experience in a smaller, B2B-focused fintech environment - this role suits someone used to close commercial context and hands-on delivery, rather than large enterprise-scale environments with extensive legacy system sprawl
  • Insurance, financial services, pricing, or marketing analytics background
  • Familiarity with UK GDPR, data residency, and regional governance
  • dbt or analytics engineering experience

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:27 min

Managing traffic and tracking costs with Databricks Unity Catalog

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

1:32 min

Recognizing the persistence and utility of spreadsheet applications

John Bettiol · World Congress 2022

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:10 min

Transitioning agile recruiting teams away from manual spreadsheet management

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

2:50 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

Videos

See all

Related articles

See all