Scientific Data Engineer (Institutional Informatics Team - Joint Genome Institut

Lawrence Berkeley National Laboratory
Berkeley, CA, United States
3 days ago
Apply on jobs.localjobnetwork.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$117,132.0 - $146,400.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Airflow Data Analysis Software Applications Big Data Databases Learning Management Systems Information Engineering Data Security Data Sharing Data Systems
+20 more
Data Warehousing Relational Databases Information Lifecycle Management Interoperability Python (Programming Language) Laboratory Information Management Systems Meta-Data Management Operational Data Store Performance Tuning Systems Architecture Systems Integration Strategies of Testing Workflow Management Systems Data Storage Technologies Concurrency Information Technology Apache Kafka Data Lakehouse Programming Languages Data Generation

Job description

Berkeley Lab’s (LBNL) Joint Genome Institute (JGI) has an opening for a Scientific Data Engineer to join the Institutional Informatics Team!

In this exciting role, you will provide technical expertise supporting raw scientific data generation and compute systems that enable laboratory operations, data analysis, and project management activities. This role will implement complex scientific and operational functional specifications requirements for automated systems. You will support the implementation and continuous improvement of core systems, including laboratory workflow orchestration, genomic data generation, metadata management, status tracking, the Laboratory Information Management System (LIMS), and the Data Warehouse/Data Lakehouse. You will be responsible for technical prowess, demonstrating good technical judgment in selecting methods and techniques for obtaining solutions, and working with the team to resolve a wide range of technical issues in creative ways. You will proactively identify technical and system needs, help refine implementation approaches, to deliver operationally and scientifically valuable solutions to the users.

The JGI is a global leader in genome science, helping shape the future of biological discovery through advanced genomic capabilities, expert support, and large-scale, AI-ready data resources. As a Department of Energy (DOE) Office of Science user facility supported by the Biological and Environmental Research (BER) program, JGI advances BER’s mission to achieve a predictive understanding of complex biological, Earth, and environmental systems in support of the nation’s energy and infrastructure security. Through world-class capabilities in genome sequencing, synthesis, transcriptomics, metabolomics, natural products, and data science, JGI supports cutting-edge research on plants, fungi, algae, microorganisms, and microbiomes. JGI is headquartered in Berkeley, CA at Berkeley Lab’s Integrative Genomics Building (IGB).

This position will be hired at the Staff or Senior level and has an anticipated start date of October 1, 2026.

We’re here for the same mission, to bring science solutions to the world. Join our team and YOU will play a supporting role in our goal to address global challenges! Have a high level of impact and work for an organization associated with 17 Nobel Prizes!, * Translate functional specifications, system designs, into implementation, system enhancements, integrations, and the evolution of shared data platforms.

  • Develop, test, deploy, maintain, and support core automated systems, services, APIs, and workflows-including the Laboratory Management System (LMS), Data Warehouse/Data Lakehouse, and Proposal/Project Management systems-to enable genomic data generation, metadata management, and JGI operations.
  • Resolve a wide range of technical issues in creative ways, integration gaps, and opportunities to optimize system architecture and operational efficiency.
  • Drive implementation efforts and ensure shared production systems meet high standards for reliability, scalability, interoperability, and performance.
  • Participate in peer technical review, documentation, and conduct continuous optimization of development workflows and service operations.

In addition to the above, the Senior Scientific Data Engineer will:

  • Translate complex scientific, operational, and user requirements into functional specifications, system designs, and implementation plans supporting system enhancements, integrations, and the evolution of shared data platforms.
  • Design, develop, test, deploy, maintain, and support core automated systems, services, APIs, and workflows-including the Laboratory Management System (LMS), Data Warehouse/Data Lakehouse, and Proposal/Project Management systems-to enable genomic data generation, metadata management, and JGI operations.
  • Proactively identify and resolve complex technical issues, integration gaps, and opportunities to optimize system architecture and operational efficiency.
  • Establish and champion engineering best practices and provide technical guidance to team members.

Requirements

  • A Bachelor’s Degree (or equivalent knowledge/training) in Computer Science or a related field and a minimum of 5 years of related professional work experience developing, integrating, deploying, and supporting production software applications and data systems that enable metadata management, workflow orchestration, data lifecycle operations, and broad user data access to scientific and operational data or an equivalent combination of education and professional experience.
  • Experience working with various database and data storage technologies including relational databases, object storage platforms, and systems supporting structured, semi-structured, and large-scale datasets.
  • Experience with data engineering and event-driven technologies such as Airflow, Kafka, or related tools.
  • Experience using AI-assisted development tools, with demonstrated sound judgment in evaluating and validating generated code for production suitability.
  • Strong working knowledge of software and data engineering fundamentals supporting large-scale production systems, including system design, APIs, testing methodologies, concurrency, reliability, scalability, interoperability, and performance optimization.
  • Proficiency in Python and experience with one or more additional programming languages.
  • Excellent communication skills, including experience organizing and presenting complex technical information to internal teams and stakeholders.
  • Demonstrated experience collaborating with stakeholders to understand project goals and translate complex scientific, operational, and user requirements into automated systems, technical specifications, and implementation plans., * A Bachelor’s Degree (or equivalent knowledge/training) in Computer Science or a related field and a minimum of 8 years of related professional work experience developing, integrating, deploying, and supporting production software applications and data systems that enable metadata management, workflow orchestration, data lifecycle operations, and broad user data access to scientific and operational data or an equivalent combination of education and professional experience.
  • Demonstrated experience provide technical leadership in shaping system architecture and technical direction across cross-functional engineering groups
  • Demonstrated experience providing technical guidance and mentorship to team members while promoting and championing engineering best practices, documentation, and continuous optimization of development workflows and operational processes., * A Master’s Degree (or equivalent knowledge/training) in Computer Science or a related field.
  • Experience working with Laboratory Information Management System (LIMS) software.
  • Experience working with laboratory operations staff and supporting laboratory workflows.

Benefits & conditions

We invest in our employees by offering a total rewards package you can count on:

  • Exceptional health benefits.
  • Generous paid time off, sick time off, and holidays.
  • A culture where you’ll belong - we are invested in our teams!, * Appointment Type: This is a full time, exempt from overtime pay (monthly paid), 2 year (benefits eligible), Term appointment with the possibility of extension or conversion to Career appointment based upon satisfactory job performance, continuing availability of funds, and ongoing operational needs.
  • Salary Range:
  • The Staff Scientific Data Engineer position has a budgeted salary range of $117,132 - $146,400 annually for job code C71.2
  • The Senior Scientific Data Engineer position has a budgeted salary range of $139,440 - $174,312 annually for job code C71.3.

It is not typical for an individual to be offered a salary at or near the top of the range for either level of the position. Salary will be commensurate with the final candidate’s qualification and experience, including skills, knowledge, relevant education, certifications, and aligned with the internal peer group.

  • Background Check: This position is subject to a background check. Any convictions will be evaluated to determine if they directly relate to the responsibilities and requirements of the position. Having a conviction history will not automatically disqualify an applicant from being considered for employment.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.localjobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

3:30 min

Approaching data problems with an engineering and strategy mindset

Becky Gandillon · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all