Data Engineer, Data Lake

Rapidhire
Charing Cross, United Kingdom
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Compensation
£ 41K

Job location

Charing Cross, United Kingdom

Tech stack

Java
Airflow
Amazon Web Services (AWS)
Apache HTTP Server
Azure
Business Intelligence
Cloud Computing
Continuous Integration
Data Architecture
Data Governance
Data Infrastructure
ETL
Data Warehousing
Dimensional Modeling
Python
Machine Learning
Meta-Data Management
Software Tools
SQL Databases
Unstructured Data
Azure
Parquet
Google Cloud Platform
Azure
Spark
Infrastructure as Code (IaC)
GIT
Containerization
Data Lake
Kubernetes
Data Lineage
Deployment Automation
Amazon Web Services (AWS)
Kafka
Data Management
Video Streaming
Stream Processing
Software Version Control
Data Pipelines
Docker
Databricks

Job description

Note: Visa sponsorship is not available for this position. Candidates must have valid work authorization in the hiring country.

  • Design, build, and maintain scalable data pipelines for batch and real-time data processing.
  • Develop, optimize, and manage enterprise Data Lake solutions on cloud or on-premises platforms.
  • Ingest, transform, and integrate structured, semi-structured, and unstructured data from multiple sources.
  • Implement ETL/ELT workflows using modern data engineering tools and frameworks.
  • Ensure data quality, governance, security, and compliance across the Data Lake environment.
  • Design and optimize data models to support analytics, reporting, and machine learning workloads.
  • Monitor, troubleshoot, and improve the performance and reliability of data pipelines.
  • Collaborate with Data Scientists, BI Developers, Analysts, and Software Engineers to understand data requirements.
  • Implement metadata management, data cataloging, and lineage tracking.
  • Automate deployment and infrastructure using CI/CD pipelines and Infrastructure as Code (IaC) where applicable.
  • Optimize storage formats (e.g., Parquet, Delta Lake, Iceberg, ORC) and partitioning strategies for performance and cost.
  • Work with cloud-native data services such as AWS, Azure, or Google Cloud data platforms.
  • Maintain documentation for data architecture, pipelines, and operational procedures.

Requirements

  • Strong experience with Data Lake architectures and modern data platforms.

  • Proficiency in SQL and Python (or Scala/Java).

  • Experience with ETL/ELT frameworks such as Apache Spark, Databricks, Azure Data Factory, AWS Glue, Apache Airflow, or Informatica.

  • Experience with cloud platforms (AWS, Azure, or GCP).

  • Knowledge of data warehousing concepts and dimensional modeling.

  • Familiarity with version control (Git) and CI/CD practices.

  • Understanding of data governance, security, and access control.

  • Experience with Delta Lake, Apache Iceberg, or Apache Hudi.

  • Knowledge of streaming technologies such as Kafka, Kinesis, or Event Hubs.

  • Experience with containerization (Docker, Kubernetes).

  • Cloud certifications (AWS, Azure, or GCP) are a plus.

Apply for this position