Senior Data Engineer

Xenon Corporation
Bloomington, IN, United States
5 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Airflow Amazon Web Services Computing Platforms Microsoft Azure Batch Processing Health Informatics Clinical Data Repository Data Architecture Data Infrastructure Data Integration Extract Transform Load (ETL)
+23 more
Supervisory Control and Data Acquisition (SCADA) Python (Programming Language) Laboratory Information Management Systems Machine Learning Message Broker Software Engineering SQL Databases Data Streaming Enterprise Data Management Snowflake Apache Spark Pyspark Information Technology Process Control Systems Pure Data Enterprise Integration AWS Data Analytics Data Management Restful APIs Data Pipelines GXP Databricks Programming Languages

Job description

We are seeking a Senior Data Engineer with extensive experience in data architecture and platform engineering to drive data initiatives for a top-tier life sciences client. This role sits at the critical intersection of scientific research informatics, clinical data systems, and manufacturing process engineering. In this position, you will own the architectural vision and hands-on execution of scalable data platforms and pipelines. You will bridge complex data domains-from small and large molecule research, genomics, and clinical trial datasets to active pharmaceutical ingredient (API) manufacturing processes, batch data, and industrial control systems. Operating in a 3-day onsite hybrid capacity in Indianapolis, you will collaborate directly with process engineers, clinical scientists, and platform teams to build high-performance data infrastructure that accelerates drug discovery and manufacturing operations., Scientific & Clinical Data Platform Architecture

  • Design, build, and maintain production-grade data pipelines and architecture tailored for scientific, clinical trial, and research informatics data (small/large molecule, genomics, proteomics, LIMS).
  • Structure complex, multi-modal clinical and scientific datasets to enable advanced analytics, enterprise reporting, and downstream machine learning models.

Process Engineering & Manufacturing Integration

  • Ingest, harmonize, and model operational technology (OT) and manufacturing process datasets, including API manufacturing pipelines, batch processing data, MES, SCADA, and OSIsoft PI systems.
  • Unify disparate laboratory and facility data pipelines into centralized, highly available enterprise data platforms.

Enterprise Data Engineering & Compliance

  • Build robust ETL/ELT pipelines using modern cloud platforms (Databricks, Snowflake, AWS/Azure), PySpark, and SQL.
  • Ensure all data pipelines and platform integrations strictly adhere to enterprise governance, data residency, and GxP regulatory standards within a heavily monitored environment.

Technical Leadership & Domain Alignment

  • Partner directly with process engineers, chemical engineering leads, and research informatics directors to translate operational friction into robust technical specifications.
  • Establish engineering best practices, data modeling standards, and pipeline monitoring frameworks across the enterprise data stack., * Not a pure Data Scientist or ML Researcher: You will not be building or training machine learning models; you are designing and scaling the underlying data architecture, pipelines, and platform infrastructure.
  • Not a non-coding Architect: This is a 100% hands-on engineering lead role requiring direct pipeline construction and technical execution.
  • Not a Fully Remote Position: This role requires a steady hybrid commitment of 3 days onsite per week at the client site in Indianapolis.

Requirements

Experience & Mindset

  • Experience: Senior-level proficiency (10-20+ years) in software development, data platform architecture, and complex ETL/ELT engineering.
  • Domain Adaptability: Demonstrated ability to engineer data pipelines across non-standard, highly specialized domains (e.g., transition between process/chemical engineering data and clinical/scientific research informatics).
  • Location & Work Auth: Must hold unrestricted US Work Authorization (no sponsorship available) and be able to work 3 days per week onsite in the Indianapolis, IN area.
  • Culture & Communication: Exceptional problem-solving mindset, strong adaptability, and the ability to articulate complex technical architecture to cross-functional engineering teams.

Must-Have Technical Stack

  • Languages & Frameworks: Advanced Python, PySpark, and expert-level SQL.
  • Data Platforms: Hands-on expertise with Databricks, Snowflake, or AWS/Azure enterprise data ecosystems.
  • Orchestration & ETL: Extensive experience with Airflow, dbt, Spark, and enterprise data orchestration engines.
  • Data Pipelines: Proven track record building streaming and batch data architectures via REST APIs, message brokers, and database integrations.

Domain Competency (Scientific & Process Focus)

  • Deep exposure to either scientific/clinical informatics (CDISC/SDTM, LIMS, clinical trials, multi-omics) OR chemical/process engineering data (API manufacturing, batch data, SCADA, MES, OSIsoft PI).

Nice-to-Haves & Certifications

  • Academic background in Chemical Engineering, Bio-process Engineering, Computer Science, or a related STEM discipline.
  • Direct experience working inside regulated GxP environments in the Life Sciences or Specialty Chemicals sectors.
  • Certifications: Databricks Certified Data Engineer Senior/Professional, Snowflake SnowPro Core/Advanced, or AWS Data Engineer Associate/Professional.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:13 min

Core components of the internal Optimize ecosystem

Dominik Schneider Dominik Schneider · World Congress 2025

2:19 min

Introduction to Apache Airflow for advanced orchestration

Alan Mazankiewicz · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

1:11 min

Deploying and running Airflow in cloud environments

Alan Mazankiewicz · LIVE

Videos

See all

Related articles

See all