> Markdown version of [/jobs/ext/1338118-senior-data-engineer](https://www.wearedevelopers.com/jobs/ext/1338118-senior-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer - **Company:** Hack The Box - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $140,000.0 - $160,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Data Analysis, BigTable, BigQuery, Software as a Service, Cloud Computing, Cloud Storage, Cyber Security, Continuous Integration, Extract Transform Load (ETL), Dimensional Modeling, Data Flow Control, Github, Python (Programming Language), Machine Learning, Online Analytical Processing, Standard Sql, Software Engineering, SQL Databases, Data Streaming, Workflow Management Systems, Google Cloud, Feature Engineering, Snowflake, Apache Spark, Build Management, Debezium, Kubernetes, Apache Flink, Production Code, Apache Kafka, Spark Streaming, Machine Learning Operations, Vertica, Restful APIs, Data Pipelines, Apache Beam, Docker - **Published:** July 18, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=ebeade119569bbf7 ## About the Role * Strong data modelling and warehouse architecture skills (dimensional modelling, event-driven, lakehouse patterns) * Hands-on experience with GCP data services - BigQuery is a must; Pub/Sub, Dataflow, Bigtable, Cloud Composer are strong pluses * Production experience with streaming pipelines on Dataflow/Beam, Flink, or Spark Structured Streaming, ingesting from Kafka and/or Pub/Sub * Solid SQL and strong Python - you write production-quality code, not just notebooks * Experience with ClickHouse or another columnar OLAP engine in production * Workflow orchestration experience with Airflow (or Prefect/Dagster) * Comfortable with dbt or equivalent transformation frameworks * Experience migrating off legacy warehouses (Snowflake, Redshift, Synapse) onto cloud-native stacks is a plus * Working knowledge of ML in production - feature engineering, feature stores, model deployment, drift monitoring, retraining * Docker & Kubernetes experience * CI/CD mindset, infrastructure-as-code sensibility, and a bias for simple, observable systems * Bonus: CDC tooling (Datastream, Debezium), Vertex AI / Feature Store ## Description Let's redefine cyber security expertise standards and connect business - community through highly engaging hacking experiences. (Find out more insights about Hack The Box culture in our career site). The core mission of the Senior Data Engineer: You will own and evolve our data pipelines on GCP - building new ones, hardening existing ones, improving data quality, and making clean, trustworthy data available across the organisation. You'll work end-to-end on streaming and batch pipelines, from CDC and event ingestion through transformation, serving, and the feature layer that powers our ML and AI products. Your day-to-day will include designing ELT/ETL processes on BigQuery and ClickHouse, building real-time pipelines on Pub/Sub and Kafka with Dataflow (and where it fits, Flink/Spark), orchestrating workflows with Airflow, and ensuring data is properly cleaned, modelled, and served for analytics, ML training, and online inference. You'll partner with ML engineers on feature pipelines, monitoring data drift, and keeping models well-fed and retrained as needed. You'll consume and build REST APIs, integrate with third-party SaaS sources, and treat infrastructure as code., You will be part of the Data, Analytics & AI team, collaborating closely with Infrastructure, Software Engineering, Product, and ML/AI engineers. We're in the middle of a GCP-native modernisation - migrating away from Snowflake toward BigQuery, Bigtable, Pub/Sub, and Dataflow - so we're looking for someone who's opinionated about clean architecture, allergic to over-engineering, and comfortable owning systems end-to-end. If retiring a legacy warehouse and standing up its replacement sounds like a good time, you'll fit right in. * ️ Technology tools & weapons you'll be using: * Cloud & warehouse: GCP, BigQuery, Bigtable, Cloud Storage * Streaming & messaging: Pub/Sub, Kafka * Processing: Dataflow (Apache Beam), with Flink/Spark where appropriate * Orchestration: Airflow (Cloud Composer) * Analytical store: ClickHouse * Languages: Python, SQL * Modelling & quality: dbt, data quality gates * Containers & CI/CD: Docker, Kubernetes, GitHub Actions / equivalent * Legacy (being retired): Snowflake The adventures that await you after becoming Senior Data Engineer at Hack The Box: * Design and build batch and streaming pipelines on Dataflow, Pub/Sub, and Kafka feeding BigQuery, Bigtable, and ClickHouse * Help drive the migration off Snowflake onto our GCP-native stack - and retire shadow pipelines along the way * Own the orchestration layer in Airflow, including SLAs, retries, and data quality gates * Model data for analytics and for ML - including feature pipelines that serve both training and low-latency online inference * Partner with ML engineers on feature stores, drift monitoring, and retraining workflows * Capture requirements from stakeholders and translate them into pragmatic, well-scoped data products * Continuously improve data quality, reliability, observability, and cost efficiency * Identify new data sources worth acquiring and integrate them cleanly, * You'll have the exhilarating opportunity to contribute to a product that is highly appreciated by users and the cybersecurity community at large * You'll experience a highly supportive and caring environment, fostering growth, flexibility, and autonomy * You'll embark on an exciting journey of continuous learning and problem-solving, leveling up as our organization grows * Most importantly, you'll have a blast at HTB because fun is an essential ingredient in our recipe for success! Just wait until you see our global meet-ups! ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)