Lead Data Engineer

Forsyth Barnes
Greater London, UK
1 day ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Adaptable Database Systems Application Programming Interfaces (APIs) Airflow Amazon Web Services Amazon S3 Application Frameworks Automation of Tests Big Data Cloud Database Continuous Integration Data Architecture Data Governance
+25 more
Data Infrastructure Data Security Data Systems Distributed Computing Environment File Transfer Protocol (FTP) SSL Extension Identity and Access Management Python (Programming Language) Machine Learning Performance Tuning SQL Databases Data Streaming User-Centered Design Management of Software Versions Workflow Management Systems File Transfer Protocol (FTP) Apache Spark Cloudformation Pyspark Infrastructure Automation Frameworks Apache Nifi Data Management Cloudwatch Terraform Data Pipelines Databricks

Job description

In this role, you will help design and evolve a modern Databricks + lakehouse architecture, enabling analytics, machine learning, and investigative teams to generate actionable insights from large-scale datasets., * Own the end-to-end design, build, optimisation, and support of scalable Spark / PySpark data pipelines (batch and streaming)

  • Define and implement lakehouse architecture standards (medallion model: bronze, silver, gold), including governance, lineage, and data quality controls
  • Design and manage secure data ingestion frameworks (e.g. Apache NiFi, APIs, SFTP/FTPS) for internal and external data sources
  • Architect and maintain secure AWS-based data infrastructure (S3, IAM, KMS, Glue, Lake Formation, Lambda, Step Functions, CloudWatch, etc.)
  • Implement orchestration using tools such as Airflow, Databricks Workflows, and Step Functions
  • Champion data quality, observability, and reliability (SLAs, monitoring, alerting, reconciliation)
  • Drive CI/CD best practices for data platforms (infrastructure as code, automated testing, versioning, environment promotion)
  • Mentor engineers on distributed data processing, performance optimisation, and cost efficiency
  • Collaborate with data science, product, and compliance teams to translate requirements into scalable data solutions

Requirements

  • Strong experience as a Senior or Lead Data Engineer with ownership of end-to-end data solutions
  • Expertise in Databricks, PySpark / Spark, SQL, and Python
  • Proven experience building and optimising large-scale data pipelines in production environments
  • Strong knowledge of cloud data architectures, particularly within AWS
  • Experience designing scalable data models and reusable frameworks
  • Hands-on experience with orchestration tools such as Airflow or similar
  • Solid understanding of data governance, lineage, and compliance requirements
  • Experience with CI/CD pipelines and infrastructure as code (e.g. Terraform, CloudFormation)
  • Strong communication skills with the ability to collaborate across technical and non-technical teams, * A hands-on technical leader who can design, build, and deliver solutions independently
  • Someone comfortable working with high-volume, high-throughput data systems
  • Strong problem-solving skills and a pragmatic, delivery-focused mindset
  • Experience mentoring engineers and setting engineering standards and best practices
  • Ability to balance technical excellence with delivery timelines

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all