Clinical Data Engineer

Bliss Clinical Research
Bedford, United Kingdom
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Intermediate
Compensation
£ 60K

Job location

Bedford, United Kingdom

Tech stack

Java
API
Artificial Intelligence
Amazon Web Services (AWS)
Data analysis
Bash
Big Data
Health Informatics
Biometrics
Clinical Data Repository
Databases
Data Validation
Information Engineering
Data Integration
Data Integrity
ETL
Data Security
Data Warehousing
Relational Databases
Cursor (Graphical User Interface Elements)
Database Design
Database Queries
Software Debugging
Digital Architecture
Firmware
Hadoop
Monitoring of Systems
Hive
Python
Shell
Machine Learning
Microsoft SQL Server
NumPy
Oracle Applications
Cloud Services
Shell Script
Signal Processing
SQL Databases
Management of Software Versions
Supervised Learning
Data Processing
Scripting (Bash/Python/Go/Ruby)
Data Storage Technologies
Informatica Powercenter
Snowflake
Spark
Pandas
Vba Programming Language
Build Tools
Data Management
Software Coding
GPT
Data Pipelines
Programming Languages

Job description

We are seeking a highly skilled Clinical Data Engineer to join our dynamic team. The successful candidate will be responsible for designing, developing, and maintaining robust data pipelines and architectures to support clinical research and healthcare data analysis. This role offers an exciting opportunity to work with cutting-edge big data technologies and contribute to innovative health solutions. The ideal applicant will possess strong technical expertise in Big Data, database management, data warehousing, and scripting, with a keen eye for detail and analysis.

You will create monitoring tools for tracking live data out in the field that can alert the research associates to any issues, work closely with ML/AI to ensure that incoming data are stored in formats that are easily ingestible and clearly labeled, and work to align datasets with multiple incoming sources of data for analysis by our team., * Develop, implement, and optimise scalable data pipelines using tools such as Apache Spark, Hadoop, and Informatica.

  • Design and maintain efficient data warehouses and databases, including Oracle, Microsoft SQL Server, and other relational systems.
  • Utilise programming languages such as Java, Python, VBA, Bash (Unix shell), and Shell Scripting to automate data processing tasks.
  • Manage cloud-based data storage solutions on AWS to ensure secure and reliable data access.
  • Collaborate with clinical teams to understand data requirements and translate them into technical solutions.
  • Perform data validation, cleansing, and transformation processes to ensure high-quality datasets for analysis.
  • Monitor system performance and troubleshoot issues related to data pipelines or storage systems.
  • Document architecture designs, workflows, and procedures for compliance and knowledge sharing purposes.
  • Stay abreast of emerging technologies in big data analytics and healthcare informatics to recommend improvements.

Data Engineering & Infrastructure

  • Build and maintain scalable ETL pipelines using Python, SQL, and APIs to ingest and process large-scale biometric and sensor data
  • Design data models and workflows that support clinical studies, internal tools, and downstream analytics
  • Manage data storage, retrieval, and archival systems in AWS, including handling long-term access and data restore workflows
  • Ensure data integrity, reproducibility, and proper versioning across evolving datasets and analyses
  • Leverage AI-assisted tools to accelerate data analysis, debugging, and code development, improving iteration speed and reducing manual effort

Clinical Analytics & Algorithm Validation

  • Analyze sleep, physiological, and behavioral datasets to evaluate product performance and validate new features
  • Perform statistical analyses (e.g., correlation, error metrics, bootstrapping, validation frameworks) to assess algorithm accuracy and clinical outcomes
  • Develop evaluation pipelines for metrics like HR/HRV accuracy, presence detection, and sleep staging
  • Build tools and structured datasets to support training and validation of machine learning models, integrating multiple data sources for supervised learning
  • Investigate edge cases, sensor issues, and data anomalies to improve model robustness

Internal Tooling & Visualization

  • Maintain and extend Python-based applications for visualizing and annotating biometric data
  • Develop interactive tools for researchers and engineers to inspect sessions, validate signals, and debug algorithms
  • Streamline workflows for clinical teams to reduce manual effort and improve reproducibility

Cross-Functional Collaboration & Communication

  • Partner with Machine Learning, Hardware, Firmware, and Product teams to build algorithms and test prototypes
  • Work with Growth and Product teams to explore user behavior and inform feature development
  • Synthesize findings into reports, dashboards, and presentations for internal teams and external audiences
  • Contribute to abstracts, posters, and conference presentations; communicate uncertainty, methodology, and tradeoffs clearly to guide decision-making

Requirements

  • 2+ years of data engineering experience with health/physiology data in a research context - you've built ETL pipelines around messy, real-world biometric or sensor datasets, not just clean CSVs
  • Advanced Python and SQL proficiency - Pandas, NumPy, time-series analysis, and production-quality scripting are daily tools, not occasional ones
  • Intermediate-to-advanced signal processing and biometric data experience - you've worked directly with heart rate, HRV, sleep staging, or similar physiological signals from wearable or embedded sensors
  • Intermediate-to-advanced statistical modeling and validation skills - you can design and execute correlation analyses, error metrics, bootstrapping, and validation frameworks independently
  • Working proficiency with AWS and Snowflake - you've built or maintained cloud-based data storage, retrieval, and archival systems, not just queried them

Bonus Points

  • Experience with clinical or regulatory trial data, familiarity with GCP/ICH guidelines.
  • Background in ML model validation or building structured training datasets for supervised learning
  • Fluency with AI-assisted development tools (Claude, Cursor, ChatGPT, Copilot) as part of your daily workflow
  • Domain knowledge in sleep science, biometrics, or wearable/embedded sensor data
  • Experience integrating internal and third-party APIs into unified data pipelines
  • Strong cross-functional communication skills - ability to translate complex analyses into clear insights for non-technical stakeholders

Skills

  • Extensive experience with AWS cloud services for scalable data management.
  • Proficiency in Java, Python, VBA, Bash (Unix shell), and Shell Scripting for automation tasks.
  • Strong understanding of big data frameworks such as Hadoop, Apache Hive, Spark, and related ecosystems.
  • Hands-on experience with relational databases, including Oracle and Microsoft SQL Server; expertise in database design is essential.
  • Knowledge of data warehousing concepts and tools to support large-scale clinical datasets.
  • Familiarity with Informatica or similar ETL tools for efficient data integration workflows.
  • Excellent analysis skills with the ability to interpret complex datasets accurately.
  • Experience working within clinical or healthcare environments is desirable but not essential.
  • Strong organisational skills with the ability to manage multiple projects simultaneously while maintaining attention to detail.

Benefits & conditions

Pulled from the full job description

  • Free parking
  • Company pension
  • On-site parking

Apply for this position