> Markdown version of [/jobs/ext/2797051-senior-data-engineer](https://www.wearedevelopers.com/jobs/ext/2797051-senior-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer - **Company:** Hack The Box - **Location:** Greater London, UK (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, BigTable, BigQuery, Continuous Integration, Data Flow Control, Python (Programming Language), Machine Learning, Online Analytical Processing, SQL Databases, Data Streaming, Workflow Management Systems, Feature Engineering, Snowflake, Build Management, Debezium, Apache Flink, Production Code, Apache Kafka, Spark Streaming, Machine Learning Operations, Vertica, Data Pipelines, Docker - **Published:** September 8, 2026 - **Apply:** https://www.collegerecruiter.com/job/2840672660-senior-data-engineer ## About the Role * Strong data modelling and warehouse architecture skills (dimensional modelling, event-driven, lakehouse patterns). * Hands-on experience with GCP data services-BigQuery is a must; Pub/Sub, Dataflow, Bigtable, Cloud Composer are strong pluses. * Production experience with streaming pipelines on Dataflow/Beam, Flink, or Spark Structured Streaming, ingesting from Kafka and/or Pub/Sub. * Solid SQL and strong Python-production-quality code. * Experience with ClickHouse or another columnar OLAP engine in production. * Workflow orchestration experience with Airflow (or Prefect/Dagster). * Comfortable with dbt or equivalent transformation frameworks. * Experience migrating off legacy warehouses (Snowflake, Redshift, Synapse) onto cloud-native stacks. * Working knowledge of ML in production-feature engineering, feature stores, model deployment, drift monitoring, retraining. * Docker and Kubernetes experience. * CI/CD mindset, infrastructure-as-code sensibility and a bias for simple, observable systems. * Bonus: CDC tooling (Datastream, Debezium), Vertex AI / Feature Store. ## Description You will own and evolve our data pipelines on GCP-building new ones, hardening existing ones, improving data quality, and making clean, trustworthy data available across the organisation. You will work end-to-end on streaming and batch pipelines, from CDC and event ingestion through transformation, serving, and the feature layer that powers our ML and AI products., * Design and build batch and streaming pipelines on Dataflow, Pub/Sub, and Kafka feeding BigQuery, Bigtable, and ClickHouse. * Help drive the migration off Snowflake onto our GCP native stack and retire legacy pipelines. * Own the orchestration layer in Airflow, including SLAs, retries, and data quality gates. * Model data for analytics and for ML-including feature pipelines that serve both training and low-latency online inference. * Partner with ML engineers on feature stores, drift monitoring and retraining workflows. * Capture requirements from stakeholders and translate them into well-scoped data products. * Continuously improve data quality, reliability, observability and cost efficiency. * Identify new data sources worth acquiring and integrate them cleanly. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)