> Markdown version of [/jobs/ext/349557-aws-databricks-engineer](https://www.wearedevelopers.com/jobs/ext/349557-aws-databricks-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AWS Databricks Engineer - **Company:** Capgemini - **Location:** London, UK (Remote available) - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Business Analytics Applications, Application Integration Architecture, Microsoft Azure, Computer Programming, Databases, Continuous Integration, Data Architecture, Data Validation, Information Engineering, Data Governance, Data Infrastructure, Extract Transform Load (ETL), Data Mapping, Data Mart, Data Warehousing, Database Queries, DevOps, Apache Hadoop, Apache Hive, Python (Programming Language), Query Optimization, Power BI, Azure Data Lake, SQL Databases, Workflow Management Systems, Data Logging, Data Ingestion, Azure Data Factory, Apache Spark, Data Layers, Amazon Relational Database Service, Data Lakes, Pyspark, Semi-structured Data, Data Management, Cloud Migration, Api Design, Cloudwatch, Azure Synapse Analytics, Data Pipelines, Databricks - **Published:** June 21, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=c45c766b74d77129 ## About the Role Do you have experience in Spark? ## Description We are looking for an experienced AWS Databricks Engineer with strong hands-on expertise in Databricks, AWS cloud services, PySpark, Spark SQL, Delta Lake, Python, SQL, and data engineering. The candidate will be responsible for designing, developing, optimizing, and supporting scalable data pipelines and Lakehouse solutions for a banking client. The ideal candidate should have strong experience in building enterprise-grade data platforms, processing large volumes of structured and semi-structured data, and implementing secure, reliable, and high-performance data pipelines in AWS-based environments. Hybrid working: The places that you work from day to day will vary according to your role, your needs, and those of the business; it will be a blend of Company offices, client sites, and your home; noting that you will be unable to work at home 100% of the time. Your Role: * Define end-to-end data architecture for Azure-based data platforms using Azure Databricks, Azure Data Factory, ADLS Gen2, Delta Lake, Azure Synapse, and Power BI. * Design scalable and secure lakehouse architecture using bronze, silver, and gold data layers. * Lead architecture and design for data ingestion, transformation, curation, data marts, reporting, and analytics solutions. * Create high-level and low-level data architecture documents, data flow diagrams, integration architecture, and data platform blueprints. * Define architecture patterns for batch, incremental, real-time, and API-based data ingestion. * Design reusable data ingestion and transformation frameworks using ADF and Databricks. * Define data models for London Market Insurance data including policy, claims, premium, broker, bordereaux, delegated authority, reinsurance, exposure, and regulatory reporting data. * Work with business analysts and insurance SMEs to understand London Market business processes and translate requirements into data architecture. * Define standards for data modelling, source-to-target mapping, data quality, reconciliation, metadata, lineage, and auditability. * Design data governance, security, access control, and compliance frameworks for insurance data. * Support cloud migration, data warehouse modernisation, reporting transformation, and legacy system decommissioning initiatives. * Review technical designs, data models, ETL/ELT pipelines, and engineering implementation. * Provide architectural guidance to data engineers working on Azure Databricks, ADF, PySpark, SQL, and Delta Lake. * Collaborate with enterprise architecture, solution architecture, security, infrastructure, DevOps, and business teams. * Define CI/CD, DevOps, deployment, monitoring, and operational support architecture for data platforms. * Identify performance, scalability, reliability, and cost optimisation opportunities across Azure data services. * Support governance forums, architecture review boards, design authority meetings, and client stakeholder workshops. Your Skills: * Strong hands-on experience with Databricks on AWS. * Strong experience with Apache Spark / PySpark. * Excellent programming skills in Python. * Strong SQL skills including complex queries, joins, CTEs, window functions, and query optimization. * AWS Secrets Manager - Secure secrets and credential management. * Amazon CloudWatch - Monitoring, logging, and alerting. * AWS Step Functions - Workflow orchestration, if applicable. * Amazon RDS / Aurora / Redshift - Source or target databases, where applicable. * Develop and maintain Databricks notebooks, workflows, jobs, and libraries. * Build reusable PySpark frameworks for ingestion, transformation, and data validation. * Implement Delta Lake features such as ACID transactions, schema evolution, time travel, and optimized storage. * Design and implement bronze, silver, and gold layers using medallion architecture. * Tune Databricks clusters for performance and cost optimization. * Monitor Databricks jobs and handle failures, retries, alerts, and job dependencies. * Implement job orchestration using Databricks Workflows, Airflow, AWS Step Functions, or similar tools. * Manage secrets, environment variables, and secure connections. * Support migration from legacy Hadoop/Spark platforms to Databricks on AWS, if required. 'We are a Disability Confident Employer: Capgemini is proud to be a Disability Confident Employer (Level 2) under the UK Government's Disability Confident scheme. As part of our commitment to inclusive recruitment, we will offer an interview to all candidates who: * Declare they have a disability, and * Meet the minimum essential criteria for the role. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)