Associate Bioinformatics Data Scientist

Signature Science, LLC
Charlottesville, VA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Data Analysis Microsoft Azure Big Data Bioinformatics Unix Computer Programming Databases Relational Databases Linux Perl (Programming Language) R (Programming Language)
+11 more
Python (Programming Language) Open Source Technology Scripting Git Containerization Information Technology Data Analytics Free and Open-Source Software Software Version Control Data Pipelines Docker

Job description

A bioinformatics data scientist is responsible for providing experimental design consulting and data analysis for large, high-throughput genomic experiments, with a focus on forensics and metagenomics. The bioinformatics data scientist will be responsible for designing and implementing annotated code for managing, manipulating, and analyzing large-scale genomic data, and for preparing thorough documentation and reporting., * Develop tools for management, analysis and interpretation of high-density microarray and whole genome sequencing data.

  • Manage, manipulate, and analyze data using a combination of R, python, and UNIX tools.
  • Use established domain-specific open-source software and tools to manipulate and analyze genomic data.
  • Implement and execute data processing workflows and automated analytic pipelines.
    • Apply literateprogramming methods to develop reproducible workflows that produce consistent, standardized tables and figures.
  • Conduct workflow benchmarking and documentation, identifying inconsistencies and resolving data problems.
  • Prepare SOPs, document source code/workflows, and write reports to summarize computational requirements, processing status, and customized analysis results.

Requirements

  • Advanced proficiency working in a Unix/Linux environment.
  • Advanced proficiency with open-source software, tools, and databases for analyzing next-generation sequencing data (whole-genome sequencing, RNA-seq, epigenetics, microbiome, and metagenomics).
  • Proficiency working with and developing using Docker and/or Singularity container technology.
  • Proficiency using version Control software (e.g., Git or similar) to manage programming code.
  • Proficiency with Python, Perl, or another scripting language.
  • Proficiency with R, RMarkdown, and the “tidyverse” tools for data analysis.
  • Preferred: Experience with NextFlow, SnakeMake, or similar workflow/pipeline management systems.
  • Preferred: Familiarity with developing and querying relational databases.
  • Preferred: Familiarity with AWS and/or Azure cloud computing.

Education/Experience:

  • BA or BS in Computer Science, Bioinformatics, or related field
  • Experience managing and analyzing large-scale datasets produced sequencing platforms and delivering solutions for managing, visualizing, analyzing, and interpreting genomic data
  • Experience using Linux/Unix text processing tools, R, and other open-source tooling to manipulate and format data, to assess data quality, and analyze data.

Clearance:

  • This position requires that the candidate be willing and able to complete a successful background screening for a security clearance. Candidates with a current security clearance will receive preference.

Supervisory Responsibilities:

  • May serve as a bioinformatics task lead.

Working Conditions/ Equipment:

  • Ability to work in varying conditions to include: traditional office environments with sedentary extended periods required for code development and testing.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:03 min

Microsoft integrating native Unix coreutils into Windows environments

Chris Heilmann +2 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

Videos

See all

Related articles

See all