Microbiologist IV (Genomic Data Engineer)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+16 more
Job description
The Microbiologist IV (Genomic Data Engineer) will provide scientific support to achieve the mission of the Coronavirus and Other Respiratory Viruses Division (CORVD). The role supports pathogen genomics, public health surveillance, outbreak detection, and epidemiological investigations through advanced genomic data engineering, integration, and analytics. The role also collaborates with multidisciplinary scientific teams, maintains technical documentation, prepares reports and scientific communications, and contributes to continuous improvement of data engineering and data management practices in support of public health objectives., * Develop, maintain, and optimize distributed data pipelines using Hadoop ecosystem tools (Hadoop Distributed File System, Spark, Hive, Impala).
- Manage large-scale ETL workflows involving genomic, epidemiological, and laboratory datasets to support bioinformatic workflows.
- Implement and optimize data validation, transformation, harmonization, and standardization workflows to ensure consistent, high-quality outputs.
- Ingest, harmonize, and manage genomic datasets from external repositories (e.g., NCBI GenBank, Sequence Read Archive) and maintain pipelines for routine updates and submissions.
- Work with genomic sequence files and associated metadata and integrate them into epidemiological and laboratory surveillance systems.
- Ensure appropriate handling of sensitive public health data and compliance with data governance expectations.
- Maintain reproducible workflows and version-controlled pipelines (e.g., Git) and prepare associated technical documentation.
- Collaborate with bioinformaticians, laboratory scientists, and epidemiologists to translate scientific questions into scalable engineered data workflows.
- Support development of analytical methods for outbreak detection and situational awareness, including Spark/SQL-based analysis.
- Document advanced data lineage, governance processes, or other high-level data management structures beyond required quality controls.
- Prepare reports, summaries, or scientific communication materials, and contribute to publications when appropriate.
- Be proficient in common programming or scripting languages, such as Python, Rust, Scala, and/or Bash
- Be present on site and attend weekly team meetings and provide updates on data engineering activities, pipeline performance, and ongoing tasks.
Requirements
- Bachelor’s degree in Bioinformatics, Data Science, Genomics, Computational Biology or a related field.
- Master’s degree is preferred in a relevant technical or scientific discipline.
Required Skils/Qualifications:
- Proficiency with Hadoop ecosystem technologies, including: Hadoop Distributed File System (HDFS), Apache Spark, Apache Hive, Apache Impala,
- Strong experience in data engineering, ETL development, and large-scale data integration.
- Experience with genomic, laboratory, epidemiological, or public health datasets.
- Ability to develop and optimize data validation, transformation, harmonization, and standardization processes.
- Experience ingesting and managing datasets from external genomic repositories such as NCBI GenBank and Sequence Read Archive (SRA).
- Proficiency working with genomic sequence files and associated metadata.
- Experience with version control systems, particularly Git.
- Knowledge of data governance, data quality management, and secure handling of sensitive health-related information.
- Proficiency in one or more programming and scripting languages such as: Python, Scala, Rust, Bash.
- Strong analytical, problem-solving, and technical documentation skills.
- Ability to collaborate effectively with multidisciplinary teams including bioinformaticians, epidemiologists, and laboratory scientists.
- Ability to work on-site and participate in regular team meetings and project updates.
Desirable Skills/Qualifications:
- Master’s degree or higher in Bioinformatics, Computational Biology, Computer Science, Data Science, Public Health Informatics, or a related discipline.
- Experience supporting pathogen genomics and infectious disease surveillance programs.
- Advanced experience with Spark-based analytics and large-scale distributed computing environments.
- Familiarity with bioinformatics workflows, genomic analysis pipelines, and sequence data management.
- Experience with analytical methods related to outbreak detection and situational awareness.
- Knowledge of public health surveillance systems and laboratory information management systems.
- Experience creating and maintaining data lineage documentation and enterprise data governance frameworks.
- Experience contributing to technical reports, scientific publications, or peer-reviewed research.
- Familiarity with cloud-based data platforms and modern data engineering practices.
- Strong communication skills with the ability to translate scientific and public health requirements into scalable technical solutions.
Benefits & conditions
Our team of talented individuals is what makes us successful. To support our team, we provide a balanced mix of benefits and programs. Your total rewards package includes competitive pay, benefits, and perks, flexible work-life balance, professional development opportunities, and performance and recognition programs. We offer a comprehensive benefits package that includes medical, dental, vision, life, and disability, voluntary benefit programs (critical illness, hospital, and accident), health savings and flexible spending accounts, and retirement 401K plan. One of our fundamental principles is to offer competitive health and welfare benefits to our team members, providing coverage and care for you and your family. Full-time employees working at least 30 hours a week on a regular basis are eligible to participate in our benefits and paid leave programs. We pride ourselves on our collaborative work environment and culture, which embraces our mission of providing financial and
About the company
Great Hill Solutions, LLC is part of the Seneca Nation Group (SNG) portfolio of companies. SNG is Seneca Holdings’ federal government contracting business that meets mission-critical needs of federal civilian, defense, and intelligence community customers. Our portfolio comprises multiple subsidiaries that participate in the Small Business Administration 8(a) program. To learn more about SNG, visit the website and follow us on LinkedIn .
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Data Engineer Salary UK
Data Analyst Salary in the UK
Highest Paying Tech Companies for Developers
Top-Paying Tech Jobs (with Salaries)