Genomics Data Operations Engineer
Role details
Job location
Tech stack
Job description
Data is the foundation of everything we do. Our Data Engineering team owns the end-to-end journey of the genomic datasets that power both our commercial products and our internal science - from disease-risk insights that reach real patients, to the discovery of novel therapeutic targets.
As a Data Engineer in this small, high-impact team, you'll turn large, complex and scientifically consequential datasets into trusted, analysis-ready resources. Your focus will be data transformation, harmonisation and quality control: working within established automated workflows, but bringing the statistical-genetics judgement to know when the numbers don't add up - and the curiosity to find out why.
As you grow, you'll take on broader work - helping define schemas for new data types and shaping how we expand our genomic data resource. It's a genuine intersection of data engineering, data modelling and genetics, with real scope to develop across both technical and scientific dimensions.
A Day in the Life
- Moving data in: ingesting large-scale genetic and genomic datasets - GWAS summary statistics, individual-level genotype and phenotype data - through a mix of scripting, automated pipelines and hands-on processing in a Linux environment.
- Safeguarding quality: interpreting QC metrics and diagnostic plots, investigating anomalies, and applying sound scientific judgement to resolve data-quality issues.
- Getting the detail right: configuring workflow parameters and curating dataset metadata so every dataset is correctly represented, fit for scientific use, and clearly described for internal scientists and external customers.
- Collaborating: partnering with software engineers and developers in an agile setting to diagnose pipeline issues, define requirements and continuously improve our ingestion and QC processes.
- Growing: contributing to schema design for new data types and helping expand our genomic data resource alongside science, product and engineering.
Requirements
- Grounded in human statistical genetics, with hands-on experience of GWAS summary statistics and/or large-scale individual-level genotype and phenotype data.
- Proficient in Python and Unix/Linux, and able to point to real bioinformatics or data-engineering work in a research or commercial setting.
- Confident reading complex data-quality outputs, and able to use scientific judgement to investigate and resolve issues.
- Well organised - happy to plan, prioritise and deliver across competing tasks at pace.
- A strong communicator who works well across multi-disciplinary teams and just as effectively on your own.
- Educated to BSc or higher in a relevant discipline - bioinformatics, computational biology, human genetics or similar - or with equivalent experience
Benefits & conditions
Pulled from the full job description
- Annual leave
- Company pension
- Private medical insurance, * Competitive Salary: Salaries are externally benchmarked annually to ensure top-of-market compensation.
- Clear Career Path: A straightforward, open progression framework means you'll always know the path to promotion and how to achieve your next career goal.
- Continuous Learning: Including external courses and a wide library of L&D materials, because your growth is our success.
Wellbeing & Time Off
- Holiday: 25 days annual leave, plus bank holidays, plus an extra 3-day company-wide shutdown at year-end.
- Financial & Health Security: Robust benefits including a market-leading pension scheme, comprehensive private health insurance for you and your family with NO excess, critical illness, and life assurance.
- Enhanced Leave: Enhanced paid family leave to support all new parents.