Data Intelligence Engineer - Design of computational data pipelines, storage & integration

Randstad
Lake Forest, IL, United States
3 days ago
Apply on www.randstadusa.com
Prepare application

Role details

Contract type
Contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$124,800.0 - $162,240.0
Working hours
Shift work

Tech stack

Application Programming Interfaces (APIs) Amazon Elastic Compute Cloud Bioinformatics Databases Information Engineering Data Governance Data Structures Data Visualization Relational Databases Data Intelligence Python (Programming Language) Machine Learning
+11 more
Open Source Technology SQL Databases Scripting Data Ingestion Delivery Pipeline Containerization Data Management Machine Learning Operations Api Design Data Pipelines Docker

Job description

job summary: What are the top 3-5 skills, experience or education required for this position:

  1. Python, SQL and scripting

  2. Database table creation and maintenance

  3. Containerization and AWS EC2 environment (or similar) familiarity

  4. Familiar with scientific data and machine learning

Experience Level = 5-7 Years

Contractor Job Description: Data Intelligence Engineer - Design of computational data pipelines, storage & integration

Project C: Improve decision making through better data capture and integration processes

We are seeking a contractor to support the capture, storage and integration of Computational Drug Discovery generated in silico data to enable faster and more informed scientific decision making. This role will help reduce the cycle time from compound design to data visualization and analysis, with the goal of enabling near real-time feedback for novel design ideas.

The contractor will work closely with Computational Drug Discovery scientists to design and implement data structures and workflows that support model inventory, metrics and results storage.

Responsibilities

Design and create database tables to support a model inventory, including ontology, model metrics, and model results.

Enable ingestion and organization of results from a range of sources, including:

custom machine learning models

co-folding affinity prediction results

custom MPO calculations

other custom scientific calculations

FEP+ results or related physics-based scoring

Facilitate model containerization and deployment on the AID self-service platform to enable automatic API deployment.

Build data ingestion workflows for result integration, including use of staging tables to manage inserts and updates of new data.

If time permits: Develop scripts to enumerate virtual molecules using Free-Wilson or MMP transformation and compute associated predictions.

Evaluate open-source models and methods as needed to support project goals.

Required Skills and Experience

Strong experience in data engineering, scientific data management, or computational chemistry/cheminformatics environments.

Proficiency with relational database design and schema development.

Experience creating and maintaining staging, integration, and data pipelines.

Working knowledge of machine learning models, model metadata, and result tracking frameworks.

Familiarity with containerization and deployment workflows such as Docker and API-based model serving.

Experience with scripting and automation, preferably in Python.

Familiarity with cheminformatics concepts such as Free-Wilson analysis, matched molecular pairs (MMPs), and virtual molecule enumeration.

Ability to evaluate open-source tools and models for scientific use cases.

Strong collaboration and communication skills for working across scientific and technical teams.

Preferred Qualifications

Experience supporting drug discovery or CDD-related data workflows.

Familiarity with FEP+ or related computational chemistry methods.

Exposure to ontology design and scientific data standardization.

Experience with cloud or platform-based self-service deployment environments.

Ability to work independently and deliver high-quality technical solutions in a contractor setting.

location: Telecommute job type: Contract salary: $60 - 78 per hour work hours: 9am to 4pm education: Bachelors

responsibilities:

  • Strong experience in data engineering, scientific data management, or computational chemistry/cheminformatics environments.
  • Proficiency with relational database design and schema development.
  • Experience creating and maintaining staging, integration, and data pipelines.
  • Working knowledge of machine learning models, model metadata, and result tracking frameworks.
  • Familiarity with containerization and deployment workflows such as Docker and API-based model serving.
  • Experience with scripting and automation, preferably in Python.
  • Familiarity with cheminformatics concepts such as Free-Wilson analysis, matched molecular pairs (MMPs), and virtual molecule enumeration.
  • Ability to evaluate open-source tools and models for scientific use cases.
  • Strong collaboration and communication skills for working across scientific and technical teams.

Preferred Qualifications

  • Experience supporting drug discovery or CDD-related data workflows.
  • Familiarity with FEP+ or related computational chemistry methods.
  • Exposure to ontology design and scientific data standardization.
  • Experience with cloud or platform-based self-service deployment environments.
  • Ability to work independently and deliver high-quality technical solutions in a contractor setting.

qualifications: Required Skills and Experience

Strong experience in data engineering, scientific data management, or computational chemistry/cheminformatics environments.

Proficiency with relational database design and schema development.

Experience creating and maintaining staging, integration, and data pipelines.

Working knowledge of machine learning models, model metadata, and result tracking frameworks.

Familiarity with containerization and deployment workflows such as Docker and API-based model serving.

Experience with scripting and automation, preferably in Python.

Familiarity with cheminformatics concepts such as Free-Wilson analysis, matched molecular pairs (MMPs), and virtual molecule enumeration.

Ability to evaluate open-source tools and models for scientific use cases.

Strong collaboration and communication skills for working across scientific and technical teams.

Preferred Qualifications

Experience supporting drug discovery or CDD-related data workflows.

Familiarity with FEP+ or related computational chemistry methods.

Exposure to ontology design and scientific data standardization.

Experience with cloud or platform-based self-service deployment environments.

Ability to work independently and deliver high-quality technical solutions in a contractor setting.

skills: API,data ingestion,Data Intelligence,data pipelines,data structures,data visualization,machine learning models,open-source,decision making,calculations,ontology,cycle time,Drug Discovery,metrics,physics,inventory,workflows

Equal Opportunity Employer: Race, Color, Religion, Sex, Sexual Orientation, Gender Identity, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status.

At Randstad Digital, we welcome people of all abilities and want to ensure that our hiring and interview process meets the needs of all applicants. If you require a reasonable accommodation to make your application or interview experience a great one, please contact HRsupport@randstadusa.com.

Pay offered to a successful candidate will be based on several factors including the candidate’s education, work experience, work location, specific job duties, certifications, etc. In addition, Randstad Digital offers a comprehensive benefits package, including: medical, prescription, dental, vision, AD&D, and life insurance offerings, short-term disability, and a 401K plan (all benefits are based on eligibility).

This posting is open for thirty (30) days.

It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

,

  • Strong experience in data engineering, scientific data management, or computational chemistry/cheminformatics environments.
  • Proficiency with relational database design and schema development.
  • Experience creating and maintaining staging, integration, and data pipelines.
  • Working knowledge of machine learning models, model metadata, and result tracking frameworks.
  • Familiarity with containerization and deployment workflows such as Docker and API-based model serving.
  • Experience with scripting and automation, preferably in Python.
  • Familiarity with cheminformatics concepts such as Free-Wilson analysis, matched molecular pairs (MMPs), and virtual molecule enumeration.
  • Ability to evaluate open-source tools and models for scientific use cases.
  • Strong collaboration and communication skills for working across scientific and technical teams.

Preferred Qualifications

  • Experience supporting drug discovery or CDD-related data workflows.
  • Familiarity with FEP+ or related computational chemistry methods.
  • Exposure to ontology design and scientific data standardization.
  • Experience with cloud or platform-based self-service deployment environments.
  • Ability to work independently and deliver high-quality technical solutions in a contractor setting.

Requirements

API,data ingestion,Data Intelligence,data pipelines,data structures,data visualization,machine learning models,open-source,decision making,calculations,ontology,cycle time,Drug Discovery,metrics,physics,inventory,workflows

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.randstadusa.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:19 min

Executing queries and scheduling pipeline jobs within DataWorks

Qiyang Duan · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

56 sec

Introduction to analytical data formats for software developers

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all