Data & AI Engineer
Role details
Job location
Tech stack
Job description
Data Ingestion Data Pipelines Bioinformatics Data Lakehouse GitHub Copilot Version Control Amazon Redshift Computer Science Data Warehousing Data Engineering Inventory Staging Data Visualization Scientific Studies Industry Standards Biological Studies Workflow Management Amazon Web Services Workflow Automation Parallel Processing Operational Databases Computational Biology Artificial Intelligence R (Programming Language) SQL (Programming Language) Engineering Design Process Extract Transform Load (ETL) Data Warehouse Architectures Python (Programming Language) Active Directory Application Mode Application Programming Interface (API), The Senior Data & AI Engineer plays a key role in driving scientific data engineering initiatives across the research pipeline, transforming complex multi-modal scientific data into reliable, actionable assets. This is a hands-on technical position focused on building modern, AI-native data platforms, onboarding multiple clinical and biological studies, and modernizing data engineering practices in a cloud environment., * Drive scientific data engineering initiatives across the research pipeline by designing and implementing robust data platforms that support clinical, biological, and multi-modal research data.
- Collaborate with IT, computational biology, and translational science leads to transform raw genomics, proteomics, and other assay data into curated, analysis-ready datasets.
- Build and maintain orchestrated ingestion pipelines for external genomics, proteomics, and other assay data sources, including source input/output, table-format writers, and row-level reconciliation.
- Develop and harden layered transformation models (staging, intermediate, and data marts) with comprehensive real-data test coverage, strong data-quality guardrails, and reusable, consolidated transformation logic.
- Implement clinical data ingestion and reconciliation workflows aligned with recognized industry standards such as SDTM and ADaM, including subject and entity resolution.
- Deliver and maintain supporting platform infrastructure, including service APIs, CI/CD pipelines, containerized deployments, observability and monitoring instrumentation, and data-warehouse performance tuning.
- Extract transformation logic and business rules from legacy analytical codebases (such as R or PySpark) and reconcile them with new platform implementations to ensure consistency and correctness.
- Translate scientific, biomarker, and bioinformatics requirements into durable data models and clearly defined, published data contracts that can be reused across teams.
- Identify repetitive or manual processes and convert them into automated workflows, guardrails, or reusable tooling, including AI-assisted workflows that accelerate future development.
- Leverage AI coding agents and tooling to build systems and workflows around AI, rather than using them only for ad hoc prompting, and ensure appropriate guardrails around AI-generated output.
- Participate in design and code reviews with an adversarial mindset, identifying edge cases, surfacing potential failure modes, and challenging suboptimal patterns to improve overall system quality.
- Contribute to onboarding multiple new scientific studies by focusing on either data ingestion or data visualization workflows, helping scale the platform to support 10+ additional studies.
- Collaborate closely with teammates to divide work effectively so that each engineer can independently own and deliver well-defined components of the data platform.
- Continuously modernize data engineering and AI practices by adopting up-to-date tools, frameworks, and methodologies in cloud-based data and AI environments., Data Ingestion Data Pipelines Bioinformatics Data Lakehouse GitHub Copilot Version Control Amazon Redshift Computer Science Data Warehousing Data Engineering Inventory Staging Data Visualization Scientific Studies Industry Standards Biological Studies Workflow Management Amazon Web Services Workflow Automation Parallel Processing Operational Databases Computational Biology Artificial Intelligence R (Programming Language) SQL (Programming Language) Engineering Design Process Extract Transform Load (ETL) Data Warehouse Architectures Python (Programming Language) Active Directory Application Mode Application Programming Interface (API) +0
Requirements
Writing Tooling Biology PySpark Research Genomics Visionary AI Agents Pipelines Automation Innovation Biomarkers Data Lakes Proteomics Code Review Scalability Data Quality Observability Data Modeling Code Coverage Reconciliation, * Demonstrated AI-native engineering practice, with hands-on experience building systems and workflows around AI coding agents (such as GitHub Copilot, Cursor, Codex, or equivalent), beyond simple prompting.
- Ability to recognize when repeated processes should be converted into automated pipelines and when AI agent output requires guardrails or additional infrastructure to ensure reliability.
- Bachelor's or master's degree in Computer Science, Data Engineering, Bioinformatics, or a closely related field.
- At least 4+ years of professional experience in data engineering with shipped production data pipelines on AWS, including services such as S3, ECS or Fargate, and Redshift or an equivalent massively parallel processing (MPP) data warehouse.
- Strong proficiency in Python for data engineering, including building data pipelines, transformations, and automation scripts.
- Strong proficiency in SQL, including writing complex queries and working with large-scale analytical databases or data warehouses.
- Working knowledge of modern data engineering libraries and frameworks used for building scalable ingestion and transformation pipelines.
- Hands-on experience with data lake and data warehouse architectures, including designing and operating data lakehouse or similar patterns.
- Familiarity with AI tooling and the ability to effectively navigate and integrate new AI tools into engineering workflows, with the understanding that this is critical to success in the role.
- Solid understanding of core data engineering principles, including data modeling, ETL/ELT design, data quality, and reliability.
- Ability to understand and work with clinical and biological (bio/omics) data, including the fundamentals of how such data is structured and used in research.
- Experience implementing clinical data ingestion and reconciliation processes aligned with standards such as SDTM and ADaM., * Experience extracting and refactoring transformation logic and business rules from legacy analytical codebases written in languages such as R or PySpark.
- Background or exposure to bioinformatics, computational biology, or related scientific domains, particularly involving genomics, proteomics, or multi-modal assay data.
- Experience designing and implementing layered transformation models (staging, intermediate, mart) with strong test coverage and data-quality checks.
- Hands-on experience delivering platform infrastructure such as service APIs, CI/CD pipelines, containerized deployments, and observability tooling.
- Familiarity with subject and entity resolution techniques in clinical data pipelines.
- Experience with adversarial or rigorous design and code review practices, including identifying edge cases and challenging design decisions to improve robustness.
- Interest in modernizing data engineering practices and staying up to date with emerging tools and methodologies in data and AI.
- Ability to work independently on a defined piece of the data platform, such as data ingestion or data visualization, while collaborating effectively with a broader team.