Data Architect Senior

University of Michigan
Ann Arbor, United States of America
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Ann Arbor, United States of America

Tech stack

Artificial Intelligence
Airflow
Bash
Bioinformatics
Clinical Data Repository
Cloud Computing
Collaborative Software
Databases
Computer Engineering
Data Architecture
Data Infrastructure
Data Systems
Software Debugging
Linux
Dicom
Python
Machine Learning
Metadata Standards
Language Modeling
Standard Sql
SQL Databases
Management of Software Versions
High Performance Computing
PyTorch
Large Language Models
Spark
Electronic Medical Records
GIT
Information Technology
Dask
Free and Open-Source Software
Data Management
Slurm
Machine Learning Operations
Data Pipelines

Job description

The Machine Learning in Neurosurgery (MLiNS) Lab at the University of Michigan is recruiting a Data Scientist / Machine Learning Engineer to build and lead the multimodal data infrastructure behind our next generation of medical AI systems. You will work with health-system data spanning radiology, digital pathology, intraoperative microscopy, and longitudinal electronic medical records, transforming complex clinical data into reliable, governed, and reusable scientific infrastructure.

This is a high-ownership role for an engineer-scientist who wants to work at the boundary of machine learning, clinical medicine, and large-scale biomedical data. Your work will directly enable foundation models, vision-language systems, and AI agents designed to improve diagnosis, surgical decision-making, and patient care.

  • Design and maintain scalable pipelines for ingesting, harmonizing, linking, and versioning multimodal clinical data across radiology, digital pathology, intraoperative imaging, and electronic medical records.
  • Develop durable SQL data models, metadata standards, cohort-building tools, and interfaces that connect imaging, pathology, clinical text, procedures, treatments, and longitudinal outcomes.
  • Create robust preprocessing pipelines for DICOM studies, volumetric MRI and CT, whole-slide pathology, and stimulated Raman histology.
  • Establish automated data-quality monitoring, validation, lineage, provenance, de-identification, and audit processes for HIPAA-regulated research environments.
  • Support distributed model training and evaluation on high-performance computing and cloud infrastructure using reproducible environments and modern MLOps practices.
  • Contribute to medical foundation models, vision-language models, clinical NLP systems, and AI agents that operate over multimodal health-system data.
  • Collaborate with clinicians, scientists, trainees, and engineers, and contribute to manuscripts, datasets, open-source software, conference presentations, and high-impact publications. Opportunities exist to lead independent technical and scientific projects

Requirements

  • Master's degree or higher in computer science, data science, computer engineering, biomedical engineering, bioinformatics, informatics, or a related field. Candidates with substantial equivalent professional experience are also encouraged to apply if permitted by the University job classification.

  • Strong Python and SQL skills, with experience building production-quality data pipelines, databases, or scientific software.

  • Fluency with Linux, Bash, Git, testing, debugging, documentation, and collaborative software-development practices.

  • Experience working with large, heterogeneous datasets and designing reliable, maintainable, and reproducible systems.

  • Experience with high-performance, distributed, or cloud computing; familiarity with SLURM is strongly valued.

  • Ability to work independently and collaborate across disciplines, with a strong commitment to scientific rigor, data stewardship, responsible AI, and clear communication.

  • Experience with DICOM, PACS, whole-slide imaging clinical data warehouses, medical imaging, computational pathology, clinical NLP, or longitudinal EHR data.

  • Experience with PyTorch and self-supervised learning, vision-language modeling, large language models, or foundation models.

  • Experience with scalable data technologies such as Spark, Dask, Ray, dbt, Airflow, Prefect, or comparable systems.

Apply for this position