IT Data Engineer - BB4352

Techdata Service Company LLC
Cambridge, MA, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Compensation
$145,600.0 - $197,600.0
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Bioinformatics Cloud Computing Cloud Database Cloud Engineering Cloud Storage Data Architecture Information Engineering
+40 more
Data Governance Data Infrastructure Extract Transform Load (ETL) Data Systems DevOps Github Apache Hive Identity and Access Management Information Lifecycle Management Python (Programming Language) Meta-Data Management DataOps Scientific Computating SQL Databases Unstructured Data Workflow Management Systems Data Logging Data Processing Enterprise Software Applications High Performance Computing Data Ingestion Cloud Monitoring Apache Spark Amazon Virtual Private Cloud (VPC) Gitlab Cloudformation Containerization Data Lakes Pyspark Kubernetes Information Technology Data Management Functional Programming Terraform Data Pipelines Api Management Serverless Computing Docker Jenkins Databricks

Job description

We are seeking a highly skilled IT Data Engineer with expertise in Databricks, Seqera Platform (Nextflow Tower), cloud data engineering, and scientific data workflows to support our Discovery R&D and data science initiatives. The ideal candidate will design, develop, and maintain scalable data platforms and automated bioinformatics/data science pipelines that enable researchers and scientists to efficiently process, analyze, and access large-scale scientific and business datasets., · Data Engineering & Platform Development

· Design, develop, and maintain scalable data pipelines using Databricks, Apache Spark, and cloud-native technologies.

· Build and optimize ETL/ELT processes for structured, semi-structured, and unstructured data.

· Develop data ingestion frameworks for research, laboratory, clinical, and external scientific datasets.

· Implement data quality, validation, monitoring, and governance processes.

· Support enterprise data lakehouse architecture and data platform modernization initiatives.

· Databricks Administration & Development

· Develop and maintain Databricks notebooks, workflows, Delta Live Tables, and Jobs.

· Create optimized Spark-based transformations and data processing solutions.

· Implement Medallion Architecture (Bronze, Silver, Gold) for data lifecycle management.

· Manage Delta Lake environments and optimize performance, scalability, and cost.

· Integrate Databricks with cloud-native services and enterprise applications.

· Seqera Platform & Scientific Workflow Management

· Deploy, configure, and support Seqera Platform (formerly Nextflow Tower).

· Develop and maintain Nextflow pipelines for bioinformatics, genomics, imaging, AI/ML, and scientific computing workloads.

· Integrate Seqera workflows with AWS cloud infrastructure and compute environments.

· Support containerized workflows using Docker and Kubernetes technologies.

· Enable reproducible, scalable, and compliant scientific data processing workflows.

· Cloud Engineering

· Design and implement cloud-based data solutions in AWS.

· Manage cloud storage solutions including S3 and data lifecycle policies.

· Develop Infrastructure-as-Code solutions using Terraform or CloudFormation.

· Implement security controls and access management following enterprise IT standards.

· Partner with data scientists, researchers, bioinformaticians, and business stakeholders to understand data requirements.

· Provide technical guidance on data engineering best practices and workflow automation.

· Troubleshoot pipeline failures, performance issues, and workflow bottlenecks.

· Contribute to platform roadmaps and continuous improvement initiatives.

· Maintain technical documentation, SOPs, and knowledge articles.

Requirements

This role requires strong experience in cloud-native architectures, data engineering best practices, workflow orchestration, and collaboration with cross-functional teams including scientists, bioinformaticians, data scientists, and IT infrastructure teams., o Bachelor’s degree in Computer Science, Information Technology, Data Engineering, Bioinformatics, or a related technical field.

o Master’s degree preferred.

o Experience

o 5+ years of experience in data engineering, cloud engineering, or analytics platform development.

o 3+ years of hands-on experience with Databricks and Apache Spark.

o 2+ years of experience with Seqera Platform (Nextflow Tower) and Nextflow workflows.

o Experience supporting scientific research, life sciences, pharmaceutical, biotech, or healthcare environments preferred.

·

o Databricks & Data Engineering, Databricks Lakehouse Platform

o Apache Spark (PySpark, Spark SQL)

o Delta Lake, Delta Live Tables (DLT)

o Databricks Workflows

o Unity Catalog

o SQL and Python

o Seqera & Scientific Computing; Seqera Platform / Nextflow Tower

o Nextflow pipeline development

o Bioinformatics workflow automation

o Docker and container technologies

o Kubernetes orchestration

o High-performance computing environments

o Cloud Technologies, AWS (required)

o S3, IAM, EC2, VPC, Lambda

o Terraform or CloudFormation

o Cloud monitoring and logging tools

o Data Technologies

o Data Lake and Lakehouse architectures

o ETL/ELT frameworks

o Data modeling

o Data cataloging and governance

o API integrations

o Data quality frameworks

o DevOps & Automation

o GitHub/GitLab

o CI/CD pipelines

o Jenkins, GitHub Actions, or similar tools

o Infrastructure as Code

o Agile and DevOps methodologies

Core Competencies

· Strong analytical and problem-solving skills.

· Excellent communication and stakeholder management abilities.

· Ability to work independently and within global cross-functional teams.

· Strong attention to detail and commitment to data quality.

· Continuous improvement mindset and passion for innovation., * Bachelor’s (Required), * Databricks: 1 year (Required)

  • SQL: 1 year (Preferred)
  • Python: 1 year (Preferred)
  • Lab/Clinical/Bio-Tech/Pharma Industry: 1 year (Required)
  • Data Engineer: 5 years (Required)
  • Seqera Platform (formerly Nextflow Tower): 1 year (Required)
  • AWS: 1 year (Required)
  • Build and optimize ETL/ELT processes : 1 year (Required)
  • Terraform: 1 year (Preferred)
  • Cloudformation: 1 year (Preferred)
  • Bioinformatics: 1 year (Preferred)
  • Databricks Lakehouse: 1 year (Preferred)

Benefits & conditions

Pulled from the full job description Health insurance Dental insurance, Job Types: Full-time, Contract

Pay: $70.00 - $95.00 per hour

Benefits:

  • Dental insurance
  • Health insurance

Application Question(s):

  • Where are you currently located?
  • What hourly rate are you seeking on a W2?

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all