Predoctoral Appointee - Machine Learning for Viral Glycosylation Prediction
Role details
Job location
Tech stack
Job description
The appointee will work closely with computational scientists, biologists, and AI researchers to develop, implement, validate, and deploy novel graph-based machine learning models capable of predicting glycosylation sites and glycan occupancy in viral proteins. The position provides an opportunity to conduct impactful research while developing expertise in machine learning, structural biology, scalable software development, and leadership within a collaborative national laboratory environment., * Design, develop, and evaluate machine learning algorithms for predicting glycosylation sites, glycosylation patterns, and glycan occupancy across diverse viral proteins.
- Develop graph neural network (GNN) architectures and evaluate alternative deep learning approaches for learning sequence- and structure-based representations of viral proteins.
- Integrate protein sequence, structural, evolutionary, and biochemical datasets to build high-quality training and benchmarking datasets.
- Design data preprocessing, feature engineering, model training, hyperparameter optimization, and benchmarking workflows for large biological datasets.
- Implement scalable software pipelines using modern machine learning frameworks (e.g., PyTorch, PyTorch Geometric, DGL, JAX, or TensorFlow) and maintain reproducible computational workflows.
- Optimize and deploy machine learning workflows on Argonne's leadership-class high-performance computing systems using distributed computing techniques where appropriate.
- Evaluate model performance using rigorous statistical analyses and compare newly developed methods against existing computational approaches.
- Collaborate closely with computational biologists, structural biologists, virologists, and computer scientists to interpret computational predictions and refine modeling strategies.
- Document software, computational workflows, datasets, and research findings to ensure reproducibility and long-term maintainability.
- Present research progress during project meetings, seminars, and laboratory reviews.
- Contribute to manuscripts, technical reports, conference presentations, and open-source software releases where appropriate.
- Participate in collaborative research initiatives across CELS and other Argonne divisions while adhering to laboratory policies and best practices for scientific software development.
- Perform additional research-related duties assigned by the supervisor that support project objectives and professional development.
Expected Outcomes
- Development of novel machine learning methods for viral glycosylation prediction.
- Implementation of scalable and reproducible computational workflows suitable for execution on leadership-class supercomputing systems.
- Validation and benchmarking of computational models against experimental and public datasets.
- Contributions to peer-reviewed publications, technical reports, and scientific presentations.
- Effective collaboration across multidisciplinary research teams.
- Development of reusable software and computational tools that support future research within CELS and the broader scientific community.
Requirements
- Recently completed Master's degree
- Demonstrated experience developing machine learning or deep learning models.
- Proficiency in Python and scientific programming.
- Experience with one or more deep learning frameworks such as PyTorch, TensorFlow, or JAX.
- Experience working with biological sequence, structural, or other scientific datasets.
- Familiarity with software engineering best practices including version control (Git), testing, and reproducible computational workflows.
- Strong analytical, problem-solving, and communication skills.
- Ability to work effectively in interdisciplinary research teams.
- Ability to model Argonne's core values of impact, safety, respect and teamwork.
Preferred Qualifications
- Experience developing graph neural networks or geometric deep learning methods.
- Experience with protein language models, structural biology, bioinformatics, or computational genomics.
- Familiarity with glycosylation biology, glycobiology, or post-translational modifications.
- Experience using high-performance computing systems, GPUs, distributed training, or parallel computing.
- Experience with scientific visualization and statistical analysis.
- Record of publications, conference presentations, or open-source software contributions.
- Familiarity with cloud computing or large-scale AI infrastructure.