Senior Data Engineer

Hack The Box
Greater London, UK
15 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Artificial Intelligence Airflow BigTable BigQuery Continuous Integration Data Flow Control Python (Programming Language) Machine Learning Online Analytical Processing SQL Databases Data Streaming Workflow Management Systems
+12 more
Feature Engineering Snowflake Build Management Debezium Apache Flink Production Code Apache Kafka Spark Streaming Machine Learning Operations Vertica Data Pipelines Docker

Job description

You will own and evolve our data pipelines on GCP-building new ones, hardening existing ones, improving data quality, and making clean, trustworthy data available across the organisation. You will work end-to-end on streaming and batch pipelines, from CDC and event ingestion through transformation, serving, and the feature layer that powers our ML and AI products., * Design and build batch and streaming pipelines on Dataflow, Pub/Sub, and Kafka feeding BigQuery, Bigtable, and ClickHouse.

  • Help drive the migration off Snowflake onto our GCP native stack and retire legacy pipelines.
  • Own the orchestration layer in Airflow, including SLAs, retries, and data quality gates.
  • Model data for analytics and for ML-including feature pipelines that serve both training and low-latency online inference.
  • Partner with ML engineers on feature stores, drift monitoring and retraining workflows.
  • Capture requirements from stakeholders and translate them into well-scoped data products.
  • Continuously improve data quality, reliability, observability and cost efficiency.
  • Identify new data sources worth acquiring and integrate them cleanly.

Requirements

  • Strong data modelling and warehouse architecture skills (dimensional modelling, event-driven, lakehouse patterns).
  • Hands-on experience with GCP data services-BigQuery is a must; Pub/Sub, Dataflow, Bigtable, Cloud Composer are strong pluses.
  • Production experience with streaming pipelines on Dataflow/Beam, Flink, or Spark Structured Streaming, ingesting from Kafka and/or Pub/Sub.
  • Solid SQL and strong Python-production-quality code.
  • Experience with ClickHouse or another columnar OLAP engine in production.
  • Workflow orchestration experience with Airflow (or Prefect/Dagster).
  • Comfortable with dbt or equivalent transformation frameworks.
  • Experience migrating off legacy warehouses (Snowflake, Redshift, Synapse) onto cloud-native stacks.
  • Working knowledge of ML in production-feature engineering, feature stores, model deployment, drift monitoring, retraining.
  • Docker and Kubernetes experience.
  • CI/CD mindset, infrastructure-as-code sensibility and a bias for simple, observable systems.
  • Bonus: CDC tooling (Datastream, Debezium), Vertex AI / Feature Store.

Benefits & conditions

  • Private health care.
  • Paid paternity leave.
  • 25 annual leave days.
  • Free lunch & snacks at the office.
  • Ticket Restaurant by Edenred.
  • Dedicated budget for training and professional development, participation in conferences.
  • Full access to our lab offerings for learning how to hack.
  • State-of-the-art equipment (Mac, iPhone, mobile plan).
  • Flexible WFH (Hybrid Model) - fully remote also available if not in Athens.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all