Senior Translational Data and AI Engineer

Kaztronix, LLC
Wilmington, DE, United States
7 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon Web Services Amazon S3 Data Analysis Apache HTTP Server Automation of Tests Bioinformatics Clinical Data Repository Code Coverage Continuous Integration
+15 more
Information Engineering Extract Transform Load (ETL) Data Warehousing Cursor (Graphical User Interface Elements) Virtual Private Networks (VPN) Python (Programming Language) Performance Tuning SQL Databases Cloudformation Containerization Pyspark Information Technology AWS Glue AWS Fargate Docker

Job description

We are seeking a contract Senior Translational Data and AI Engineer to help modernize how biomarker and clinical data are ingested, transformed, and delivered within Computational Discovery. The group is consolidating a fragmented set of ETL processes onto a modern, AWS-native lakehouse platform. Core infrastructure is already in place; this role adds the engineering capacity to bring real scientific and clinical data through production pipelines across multiple studies, and to serve it reliably to downstream analytics, AI, and visualization consumers. The successful candidate will work as a hands-on technical partner alongside the lead engineer, contributing across the full stack from source ingestion through curated delivery. This is a role that pairs strong data-quality instincts with an AI-native way of working: not just using AI coding agents, but building scalable, guarded workflows around them. Top-level architecture and design decisions remain owned by the lead engineer; we are looking for a strong senior engineer who executes independently within that direction and is comfortable in scientific data domains or ramps into them quickly., * Build and maintain orchestrated ingestion pipelines for external genomics, proteomics, and other assay data sources, including source IO, table-format writers, and row-level reconciliation.

  • Develop and harden layered transformation models (staging, intermediate, and mart) with real-data test coverage, data-quality guardrails, and reusable, consolidated logic.
  • Implement clinical data ingestion and reconciliation paths against recognized standards (e.g., SDTM, ADaM), including subject and entity resolution.
  • Deliver supporting platform infrastructure: service APIs, CI/CD pipelines, containerized deployments, observability instrumentation, and data-warehouse performance tuning.
  • Extract transformation logic and business rules from legacy analytical code (e.g., R, PySpark) and reconcile them against new platform implementations.
  • Translate scientific and biomarker requirements from research and bioinformatics partners into durable data models and published data contracts.
  • Identify repetitive processes and convert them into automated workflows, guardrails, or reusable tooling including AI-assisted workflows that make future work faster.
  • Participate in adversarial design and code reviews, identifying edge cases and pushing back on suboptimal patterns.
  • Collaborate with the lead engineer on design decisions and support delivery velocity through paired working sessions and PR reviews.
  • Ensure all work meets reproducibility standards: CI on every PR, automated tests, and no ad-hoc notebook-based production processes.

Requirements

  • AI-native engineering practice: demonstrated experience building systems and workflows around AI coding agents (Claude Code, Cursor, Codex, or equivalent) not just prompting them. You recognize when a repeated process should become an automated pipeline, when agent output needs guardrails, and when to build infrastructure that makes future work faster. Surface-level tool usage is insufficient.
  • Education: Bachelor’s or master’s degree in Computer Science, Data Engineering, Bioinformatics, or a related field.
  • Experience: 4+ years of professional experience in data engineering with shipped production pipelines on AWS (S3, ECS/Fargate, Redshift or equivalent MPP).
  • Strong proficiency in Python and SQL with working knowledge of modern data engineering libraries.
  • Solid, hands-on experience with dbt and a workflow orchestration tool (Dagster, Airflow, or Prefect).
  • Data quality instinct: track record of catching silent failures, questioning data correctness assumptions, and noticing lossy joins or incomplete deliveries.
  • Working understanding of lakehouse architecture patterns, ETL processes, and schema design for complex multi-modal datasets.
  • Comfort working with scientific, biomarker, or other complex domain data or a demonstrated ability to ramp on unfamiliar scientific domains quickly.
  • Ability to handle PHI-adjacent clinical data under contractor policy (background check, compliance training, VPN access).
  • Willingness to work within legacy codebases (R, PySpark) to extract business rules and validate new implementations.
  • Excellent communication skills and ability to work in an embedded pair model with tight feedback loops., * Experience building or maintaining tooling around AI coding agents (custom commands, subagents, evals, or guardrails) rather than only consuming them.
  • Direct experience with Apache Iceberg, AWS Glue Catalog, or lakehouse table formats.
  • Preferred fluency reading genomic data (VAF, HGVS nomenclature, VCFs, CNV/fusion semantics).
  • Familiarity with clinical data standards including SDTM, ADaM, and CDISC.
  • Pharma, clinical research, or life sciences background.
  • Experience with containerization (Docker/ECS) and infrastructure-as-code (CloudFormation).
  • Proficiency in R for interoperability with bioinformatics teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · WWC 2024

Videos

See all

Related articles

See all