> Markdown version of [/jobs/ext/2258751-data-engineer-bioinformatics](https://www.wearedevelopers.com/jobs/ext/2258751-data-engineer-bioinformatics). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer (Bioinformatics) - **Company:** Our Future Health - **Location:** London, UK (Remote available) - **Salary:** £74,000.0 - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Agile Methodology, Airflow, Apache HTTP Server, Unit Testing, Microsoft Azure, Bioinformatics, Code Review, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Transformation, Distributed Systems, Python (Programming Language), Workflow Management Systems, Parquet, Cloud Platform System, Apache Spark, Data Lakes, Kubernetes, Software Version Control, Data Pipelines, Docker, Databricks - **Published:** August 26, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5854678933 ## About the Role We're looking for a Data Engineer with a solid understanding and experience of bioinformatics, in particular tools and methods associated with genomic data. You can design, build and test pipelines using a range of different technologies. You know how to create repeatable and reusable products and can communicate to and between technical and non-technical stakeholders, with the ability to facilitate discussions and manage different perspectives within a multidisciplinary team including scientists, software engineers, product managers and other data engineers., We welcome applications from all who may not feel they match the full criteria, so if you have most of the below, we'd like to hear from you: * Experience building and maintaining robust, scalable and efficient data pipelines. Capable of processing very large amounts of data based on feeds from multiple systems using a range of different technologies. * Can listen to the needs of technical and business stakeholders and interpret them, and effectively manage stakeholder expectations. * Detailed knowledge and understanding of genomic data (experience in genotyping and imputation is advantageous). * Experience using bioinformatics file standards (VCF, BGEN etc) and tools (PLINK, bcftools, QCtools etc) * Highly proficient in Python. * Highly proficient in version control and Git/GitHub. * Experience of workflow management tools, e.g. Nextflow, WDL/Cromwell, Airflow, Prefect, Dagster * Understanding of containerisation (e.g. Docker) and deployment (e.g. Kubernetes). * Good understanding of cloud environments (ideally Azure), distributed computing and scaling workflows and pipelines * Understanding of common data transformation and storage formats, e.g. Apache Parquet. * Awareness of data standards such as GA4GH ( https://www.ga4gh.org/) and FAIR (https://www.go-fair.org/fair-principles/). * Experience with Spark, Databricks, data lakes. * Follow best practices like code reviews, clean code and unit tests. * Experience working in an agile development team. ## Description * Support the build of re-usable data pipelines used to identify prospective clinical trial participants. * Produce logic for data transformation steps as code, which meets the requirements for our end users and builds well curated, accessible and quality controlled data for analysis. * Developing prototypes for pipelines for complex transformations drawing on existing workflows developed in industry and academia. * Keep abreast of best practice in data engineering across industry, research and Government and facilitating the adoption of standards. * Providing technical input into the upstream parts of the data pipeline, including the specification and transfer of data from data providers. * Routine ad-hoc data curation activities requiring hands on development of bespoke ETL cleaning scripts using languages such as Python. * Working with researchers to understand the data requirements and work with them to deliver the data needed for their projects., You'll be joining at an important point as we scale our clinical research recruitment capabilities and move from manual scientific workflows towards reusable engineering solutions. That means the opportunity to build things that don't exist today, tackle challenging problems with huge datasets and see a clear connection between your engineering work and real-world health research. We're looking for someone pragmatic, collaborative and comfortable working through ambiguity, someone who enjoys solving difficult problems rather than simply maintaining established systems. It's a rare combination of science, engineering, scale and purpose. Hiring process We feel hiring should be transparent and give you a real sense for what Our Future Health is like. Here's what you can expect: * Initial chat with our Talent team (30 min) to get to know each other, discuss the role, and answer any questions you have. * 1st interview with our Head of Data Engineering (30 min) - this is an opportunity to learn more about each other and align on role expectations. * Technical interview with 2-3 members from our engineering team (60 min). This will include a short task you'll need to prepare in advance of the interview and is designed to get a sense for how you'd approach the real-world responsibilities of this role. * Final stage competency interview (60 min) with a small cross-functional panel of your potential new colleagues. This will focus on how you collaborate, work with stakeholders, and navigate ambiguity and challenges in real-world projects., At Our Future Health, we recognise the importance of having a diverse workforce and ensuring that all candidates, regardless of their background, have equitable access to our application process. We proactively encourage applicants who identify as having a disability, neurodiversity, or long-term health conditions to let us know if they require any reasonable adjustments as part of their application process. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Fullstack Developer Salary UK](https://www.wearedevelopers.com/magazine/251-fullstack-developer-salary-uk)