Lead Data Architect
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+27 more
Job description
- Own the technical roadmap for the platform, aligning engineering outcomes to business priorities and operational resilience.
- Lead solution design for new integrations and platform capabilities, balancing pace with long-term maintainability.
- Set engineering standards (coding, testing, documentation, observability, security) and ensure consistent adoption across the team.
- Provide hands-on contribution where needed (critical ETLs, framework components, design spikes, performance fixes).
- Design and develop robust ETL pipelines orchestrated in Apache Airflow, including scheduling strategies, dependency management, retries, backfills, and processing.
- Establish patterns for onboarding new source systems and delivering data to downstream vendor systems with strong reliability and traceability.
- Drive data quality controls (validation, reconciliation, exception handling) and operational runbooks.
- Evolve the multi-tenant architecture: tenant isolation, configuration management, metadata-driven pipelines, and scalable runtime patterns.
- Optimise for performance, cost, and reliability across environments (dev/test/prod), including capacity planning and operational SLAs.
- Improve build and deployment automation for both platform services and Airflow DAGs (versioning, packaging, promotions, approvals).
- Implement CI/CD best practices: automated testing, static analysis, security scanning, and controlled rollout/rollback strategies.
- Reduce manual toil by investing in self-service onboarding, templates, and engineering enablement tooling.
- Implement end-to-end data lineage (source * platform * vendor), including metadata capture and a visual representation that supports auditability and faster incident resolution.
- Enhance observability across pipelines and services: logging, metrics, alerting, and dashboards; drive measurable improvements in MTTR and change failure rate.
- Partner with product owners, data owners, architects, vendors, and governance teams to shape requirements and manage delivery expectations.
- Lead agile ceremonies and engineering planning; remove blockers and maintain delivery momentum.
- Mentor engineers through code reviews, design reviews, and structured coaching; foster a culture of continuous improvement.
Technologies:
- Airflow
- BigQuery
- CI/CD
- Cloud
- ETL
- GCP
- IAM
- Python
- Security
- Terraform
- Architect
More:
We are leading the engineering delivery and technical direction for a scalable, multi-tenant data integration platform on Google Cloud Platform (GCP). Our platform uses Apache Airflow to orchestrate ETL pipelines and enables reliable movement of data from source systems through the platform into downstream vendor systems. We focus on building high-quality ETLs, improving automation and deployment pipelines, and implementing end-to-end data lineage with clear, visual traceability of data movement. We are looking for a pragmatic, hands-on technical leader who enjoys getting things done, raises engineering standards without slowing delivery, and builds an inclusive team environment where different viewpoints improve outcomes.
Requirements
- Proven experience as a Technical Lead / Lead Engineer delivering data engineering or integration platforms in cloud environments.
- Strong hands-on experience building production-grade ETL/ELT pipelines and orchestration with Apache Airflow.
- Solid experience on Google Cloud Platform, such as data services (e.g., BigQuery, Cloud Storage, Pub/Sub, Dataflow) and/or container platforms (e.g., GKE, Cloud Run).
- Experience with IAM, networking, secrets management, and environment separation.
- Strong software engineering foundations: Python, APIs, configuration management, testing strategy, and performance tuning.
- CI/CD expertise (pipeline tooling, release controls, automated quality gates) and Infrastructure-as-Code experience (e.g., Terraform).
- Experience designing for reliability: retries, idempotency, dead-letter patterns, backpressure handling, and operational resilience.
- Ability to translate complex technical topics for mixed audiences and drive alignment across teams.
- Desirable: experience implementing or integrating data lineage / metadata management tooling (e.g., Data Catalog-style metadata, OpenLineage, DataHub, etc.).
- Desirable: experience with multi-tenant platforms (tenant isolation models, metadata-driven frameworks).
- Desirable: knowledge of data governance concepts (data quality, auditability, retention, access controls).
- Desirable: experience integrating with external vendor systems and managing vendor technical dependencies.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Data Engineer Salary UK
Top Big Data Technologies That You Need to Know
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production