> Markdown version of [/jobs/ext/3576015-data-engineer](https://www.wearedevelopers.com/jobs/ext/3576015-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Creditsafe - **Location:** Cardiff, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Data Infrastructure, Cursor, Distributed Computing Environment, Amazon DynamoDB, Python (Programming Language), Raw Data, Microsoft Copilot, DataOps, Perplexity AI, Software Systems, Data Processing, Apache Spark, Build Management, Low Latency, AWS Glue, AWS Fargate, Functional Programming, Data Delivery, Data Pipelines, Serverless Computing - **Published:** October 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=bdb4d2383495fb0e ## About the Role You understand and can implement data engineering best practices. You demonstrate the ability to write clean & efficient code and combine it with the cloud environment for best performance. You have a minimum 6 years of development experience within a commercial environment creating production grade data pipelines in python. You are looking to grow your skills through daily technical challenges and enjoy problem solving, whiteboarding in collaboration with team. You have excellent communication skills, and ability to explain your views clearly to the team and are open to understanding the views of others. You have a proven track record to draw from a deep and broad technical expertise to mentor engineers, complete hands-on technical work, and provide leadership on complex technology issues. You share your ideas collaboratively via wikis, discussions boards, etc and share any decisions made, for the benefit of others. You take ownership of end-to-end deliverables from planning all the way to production. ## Description Join a team of highly technically skilled engineers who are designing Creditsafe's new Consumer Bureau with data quality, lineage, and auditability as primary goals whilst maintaining high throughput and scalability. The data platform will be built using AWS infrastructure such as S3 and S3-tables(Iceberg), Aurora Serverless, DynamoDB, Lambda, Fargate, ECS, Airflow, Spark and Kinesis. The platform is expected to manage over one hundred million objects with a high volume of daily updates handling addition, deletion, and correction of our data and indexes in an auditable manner. The platform will employ a sophisticated matching algorithm to enable us to assign incoming updates to existing objects. Our data processing application is entirely based on Python and designed to efficiently transform incoming raw data volume into API consumable schema. The team is also building highly available and low latency APIs to enable our clients with faster data delivery. WHAT YOU'LL OWN Contributing actively to the codebase and participating in peer reviews. This is not a passive "take a ticket and close it" role - you will drive development from 3-Amigo's refinement sessions all the way to post-deployment observability. Design and build metadata driven, event based distributed data processing platform using technologies such as Python, Spark, Airflow, Iceberg, DynamoDB, AWS Glue, S3. Play an active role in the design, development, and deployment of our business-critical system. Building and scaling Creditsafe's data pipelines combining batch and streaming approaches and targeting sub-500ms latency serving. Understanding domain data to make recommendations to improve the existing product. Employing various AI tools such as Perplexity, Cursor, CoPilot, etc to aid in all aspects of your work., * Written design culture. Anything non-trivial gets an RFC. Decisions are recorded. We disagree in documents, not in meetings.