> Markdown version of [/jobs/ext/580202-bioinformatics-data-scientist](https://www.wearedevelopers.com/jobs/ext/580202-bioinformatics-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Bioinformatics Data Scientist - **Company:** Spectrix Analytical Services, LLC - **Location:** Cambridge, MA, United States - **Salary:** $83,200.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Business Analytics Applications, Bioinformatics, Cloud Engineering, Computational Biology, Data Structures, Decision Support Systems, Oracle Discoverer, R (Programming Language), Systems Analysis, Python (Programming Language), Machine Learning, Meta-Data Management, RStudio, Scientific Computating, Software Deployment, Software Engineering, SQL Databases, Parquet, Cloud Platform System, Flask (Web Framework), Large Language Models, Multi-Agent Systems, Prompt Engineering, Generative AI, Git, Fastapi, Containerization, Data Lakes, Information Technology, Free and Open-Source Software, Feature Selection, Machine Learning Operations, Streamlit Framework, Software Version Control, Data Pipelines, Docker - **Published:** June 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=d7716b1a79b7d737 ## About the Role Do you have experience in Translational research?, The ideal candidate is a data scientist with strong computational biology, statistical, machine learning, and software engineering foundations, with the ability to work across biological interpretation, cloud-based workflows, and emerging AI/LLM-enabled scientific capabilities., · PhD in Bioinformatics, Computational Biology, Data Science, Statistics, Computer Science, Systems Biology, Biology, Biochemistry, Chemistry, Engineering, or a related quantitative scientific field. · Hands-on experience analyzing high-dimensional biological or biomedical datasets, such as proteomics, LC-MS, Olink, SomaScan, transcriptomics, single-cell data, spatial omics, genomics, genetics, or other omics/modalities. · Strong proficiency in R and/or Python for data analysis, statistical modeling, visualization, and reproducible scientific computing. · Solid understanding of experimental design, data QC, normalization, missing-value assessment and imputation, feature selection, and advanced statistical modeling approaches for high-dimensional biological data, including linear models, mixed-effects models, Bayesian methods, biological signal deconvolution, pathway or feature-level interpretation, and communication of biological findings. · Experience working with scalable computing and cloud environments, including AWS services and workflow-based analysis systems for large scientific datasets. · Proficiency with machine learning, AI, and LLM-powered analytical systems, including the development of robust agentic frameworks or multi-agent architectures for scientific workflows. · Familiarity with drug discovery, translational research, biomarker discovery, perturbation biology, pharmacodynamic studies, disease biology, or related biomedical research contexts. · Ability to work independently on complex analytical problems and communicate results clearly to scientific stakeholders. Preferred Qualifications · Postdoctoral, industry, or equivalent applied research experience after PhD. · Direct experience with computational proteomics across one or more platforms, including LC-MS proteomics, phosphoproteomics, DIA/SWATH, DDA, TMT, label-free quantification, PTM analysis, spectral library generation and prediction, affinity-based proteomics such as Olink or SomaScan, and related analytical workflows. · Hands-on experience with proteomics software, outputs, or data structures from tools such as DIA-NN, Spectronaut, MaxQuant, FragPipe, Proteome Discoverer, Skyline, or comparable platforms. · Experience building reusable scientific workflows, analytical pipelines, applications, dashboards, APIs, reports, or self-service data products for scientific users. · Experience deploying analytical workflows or data products on AWS or comparable cloud platforms using workflow orchestration, containerization, and scalable execution frameworks such as Nextflow, Snakemake, Airflow, Docker, AWS Batch, ECS, or comparable systems. · Experience developing and deploying R and Python software and data products using technologies such as Posit/RStudio, Posit Connect, Shiny, Dash, Streamlit, FastAPI, Flask, or comparable scientific computing and application delivery platforms. · Strong software engineering practices, including Git/version control, modular code design, documentation, testing, and reproducible workflow development. · Hands-on experience developing LLM-enabled applications or workflows using Claude or other large language models, including RAG systems, tool-using agents, prompt engineering, evaluation frameworks, LLMOps concepts, or scientific knowledge extraction. · Experience applying machine learning or foundation-model approaches to biological data, including representation learning, multimodal modeling, classification/regression, embedding-based retrieval, generative AI, or related methods. · Experience with data modeling, SQL, Parquet, metadata management, data lake architectures, or large-scale biological data warehouses. · Strong publication record, open-source contributions, or demonstrated delivery of reusable computational tools, analytical platforms, scientific workflows, or production-quality internal data products. Strong Differentiators · Ability to bridge computational proteomics, biological interpretation, cloud engineering, bioinformatics methodology development, and AI/ML/LLM workflow implementation for life science applications. · Demonstrated success building tools, analytical methods, or platforms that were adopted by experimental, translational, or computational scientists. · Contributions to peer-reviewed publications in bioinformatics, computational biology, proteomics, machine learning, systems biology, or related fields. · Experience designing AI-assisted, machine-learning, or agentic workflows that are reproducible, traceable, scientifically reliable, and suitable for biological and biomedical research. · Strong understanding of how to connect omics data, pathway biology, perturbation data, genetics, and drug discovery questions into reusable analytical systems. · Ability to develop and publish novel analytical methodologies when appropriate. · Track record of applying AI and machine learning techniques to life science datasets, including biomarker discovery, target identification, predictive modeling, knowledge extraction, or multi-omics integration. · Ability to help shape future scientific AI strategy rather than only execute predefined analyses., * Doctorate (Required) ## Description · Analyze and interpret large-scale proteomics and multi-omics datasets to support biomarker discovery, pharmacodynamic analysis, pathway and causality inference, disease biology, and drug discovery programs. · Develop scalable, reproducible, and cloud-enabled analytical workflows, data pipelines, reports, dashboards, APIs, and data products for scientific users. · Apply statistical modeling, machine learning, and AI/LLM-enabled approaches to improve biological interpretation, knowledge extraction, workflow automation, and scientific decision support. · Integrate proteomics data with orthogonal modalities such as transcriptomics, genomics, genetics, perturbation data, metadata, and translational annotations. · Collaborate with computational scientists, mass spectrometry scientists, discovery biologists, translational researchers, and data science teams to define analytical strategies and communicate results clearly. · Promote best practices in data QC, reproducible analysis, workflow development, software engineering, and responsible use of AI-assisted scientific tools. ## Related Videos - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [The Fastest-Growing Tech Sectors to Look Out for in 2025](https://www.wearedevelopers.com/magazine/373-the-fastest-growing-tech-sectors-to-look-out-for-in-2025) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)