Senior Data Engineer

Auction Technology Group
London, UK
24 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Clean Code Principles Query Performance A/B Testing Airflow Data Analysis Automated Storage and Retrieval Systems Automation of Tests Cloud Storage Code Review Continuous Integration Data as a Services Data Infrastructure
+37 more
Data Sharing Software Debugging Distributed Computing Environment Elasticsearch Monitoring of Systems JSON Python (Programming Language) Machine Learning Metadata Repositories Modular Design Performance Tuning Software Engineering Data Streaming Transaction Data Privacy Controls Parquet Cloud Platform System Real Time Systems Sql Optimization Snowflake Apache Spark Model Validation Git Data Layers Containerization Kubernetes Infrastructure Automation Frameworks Apache Flink Avro Apache Kafka Spark Streaming Data Management Machine Learning Operations Software Coding Terraform Data Pipelines Docker

Job description

We are investing in the data foundation behind ATG’s customer experiences and business decisions. As a Senior Data Engineer on the Data Enablement team, you will build production-grade data products that serve analytics, search, recommendations, personalization, and machine learning. You will work closely with product managers, analysts, data scientists, ML engineers, and software engineers to turn ambiguous needs into dependable, well-documented datasets and pipelines. This is a hands-on engineering role for someone who cares about maintainability, data quality, and measurable outcomes. You will help shape standards and architecture while still writing code, reviewing designs, troubleshooting failures, and improving the platform., What you will do Build durable data products

  • Design, build, and operate batch and event-driven pipelines for auction, inventory, customer, and transaction data.
  • Develop reusable transformation models and curated datasets in Snowflake and dbt for analytics and operational use cases.
  • Orchestrate complex dependencies with Airflow, Dagster, or a comparable workflow platform.
  • Design data models and interfaces that are clear, scalable, and easy for downstream teams to use. Raise reliability and data quality

  • Define data contracts, validation rules, freshness expectations, lineage, and service-level objectives for critical datasets.
  • Implement automated testing, anomaly detection, alerting, and observability across the data lifecycle.
  • Own production issues through diagnosis, recovery, root-cause analysis, and prevention.
  • Improve query performance, warehouse efficiency, and cloud cost without compromising reliability. Enable machine learning and customer experiences

  • Create versioned training, validation, and inference datasets for search, recommendations, personalization, and other ML products.
  • Partner with ML engineers and data scientists to make feature computation reproducible and consistent across experimentation and production.
  • Support experimentation by delivering trustworthy exposure, interaction, and outcome data for A/B testing and model evaluation. Strengthen engineering practices

  • Apply software engineering practices to data work, including modular design, code review, automated testing, CI/CD, and infrastructure as code.
  • Improve documentation, discoverability, access controls, and governance for shared data products.
  • Contribute to architectural decisions, technical standards, and pragmatic platform improvements.
  • Mentor engineers and help the team make sound trade-offs among speed, scale, cost, and maintainability., * Teams can find and use well-documented data products without unnecessary handoffs.
  • New analytics and ML use cases move from idea to production faster because the underlying data is ready and reusable.
  • Data platform performance and cost improve as ATG scales.
  • Engineering standards become easier to follow and are adopted across the team.

Requirements

  • Five or more years of experience building and operating data pipelines or data platforms in production.
  • Strong Python and advanced SQL skills, including testing, debugging, performance tuning, and maintainable code design.
  • Hands-on experience with Snowflake or another modern cloud data platform, plus practical knowledge of dimensional and analytical data modeling.
  • Production experience with dbt or a comparable transformation framework and with Airflow, Dagster, Prefect, or similar orchestration tooling.
  • Experience with AWS data services and cloud storage; equivalent experience on another major cloud platform is welcome.
  • A working understanding of data quality, lineage, observability, data contracts, and operational ownership.
  • Comfort with Git, code review, automated testing, and CI/CD for data pipelines.
  • Clear communication and the ability to work across product, analytics, software engineering, data science, and ML teams.
  • A bachelor’s degree in a relevant field or equivalent practical experience. Useful, but not required

  • Event streaming or real-time processing with Kafka, Kinesis, Flink, Spark Structured Streaming, or similar technologies.
  • Distributed processing with Spark and familiarity with Parquet, Avro, JSON, and open table formats.
  • Search or retrieval systems such as Elasticsearch/OpenSearch, vector databases, or embedding pipelines.
  • Feature stores, ML data pipelines, model monitoring, or other MLOps capabilities.
  • Infrastructure as code and container platforms, including Terraform, Docker, or Kubernetes.
  • Data catalogs, semantic layers, master data management, or metadata-driven governance.
  • Experience with ecommerce, marketplaces, auctions, GDPR, privacy controls, or regulated data environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:52 min

Customizing block storage tiers and formats

Ricardo Sueiras Sueiras · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · World Congress 2025

Videos

See all

Related articles

See all