> Markdown version of [/jobs/ext/2311564-lead-data-engineer-identity](https://www.wearedevelopers.com/jobs/ext/2311564-lead-data-engineer-identity). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Engineer - Identity - **Company:** Kargo - **Location:** Greater London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Databases, Continuous Integration, Data Systems, Github, Python (Programming Language), Online Analytical Processing, Prometheus, SQL Databases, Data Streaming, Aerospike, Data Ingestion, Snowflake, Grafana, Apache Spark, Data Layers, Kubernetes, Low Latency, Apache Kafka, Vertica, User Identification - **Published:** August 30, 2026 - **Apply:** https://www.collegerecruiter.com/job/2830561532-lead-data-engineer-identity ## About the Role Techies who want to build the future. Creatives who want to design it better. Communicators to win business. Collaborators to build it. Data pros who turn numbers into insights. Product builders who turn ideas into innovations. Anyone eager to be on a team that doesn't stop to ask what's next, because they're already building it., * You've designed and owned large-scale, interdependent data systems, including at least one you built from scratch, and you turn ambiguity into a sequenced roadmap with Product and Data Partnerships. * You've led engineers, setting direction, reviewing work, developing people, while staying hands-on. * You have mastery of Python, Airflow and Spark, and write transformations that are idiomatic, testable and tuned for cost and performance; you write SQL for Snowflake with the same discipline. * You're at home in AWS and Kubernetes, can read infrastructure logs to diagnose failures and slowness, and have worked with third-party APIs inside ingestion pipelines. * You're fluent with AI tooling in your own work, and you think about what makes a codebase legible to it., * Iceberg or a comparable table format at production scale. * Streaming or near-real-time processing (Kafka, Redpanda or similar). * Low-latency stores such as Aerospike * Experience with OLAP databases like Clickhouse. ## Description * Build a new identity graph. Take stock of what we have today, set its direction, and sequence the rollout: identifier sync, translation, clustering (with Data Science), opt-out handling. * Standardize partner and client onboarding across web, CTV and mobile identifiers, including cleanroom onboarding, so each new feed costs less to stand up than the last. * Ready the identity audience data layer for self-serve: creation, activation, state, and the reporting clients will discover audiences through. * Own and raise the bar on the domain's observability and alert response. Inventory today's signals, monitors and alerts, centralize them, and bring each to standard: a freshness and quality commitment, context for AI-assisted triage, and a runbook. * Lead and grow the domain's data engineers. Define the standards for testability, cost efficiency and the patterns worth repeating, then raise the team to them through your own code, reviews, and knowledge-sharing, recording decisions in ADRs., * Identity resolution or graph work in AdTech: matching, device and household graphs. * Privacy and consent obligations: opt-outs, deletion, GDPR and CCPA. * Data cleanrooms for partner or client onboarding. * CI/CD with GitHub Actions/ArgoCD; monitoring with VictoriaMetrics/Prometheus/Grafana. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)