> Markdown version of [/jobs/ext/2696171-senior-data-scientist-biologics-discovery](https://www.wearedevelopers.com/jobs/ext/2696171-senior-data-scientist-biologics-discovery). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Scientist, Biologics Discovery - **Company:** Johnson & Johnson - **Location:** Spring House, PA, United States - **Experience:** Expert - **Salary:** $109,000.0 - $174,800.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bioinformatics, Computational Biology, Data Infrastructure, Experimental Data, Monitoring of Systems, Python (Programming Language), Machine Learning, Language Modeling, Molecular Modeling, SQL Databases, Pytorch, Scikit Learn, Information Technology, Machine Learning Operations - **Published:** September 3, 2026 - **Apply:** https://www.careerjet.com/job/us92dd304300204cff4f75b6bcb4608945/eaa ## About the Role * Master's or Ph.D. in Computer Science, Machine Learning, Computational Biology, Bioinformatics, Statistics, or a related field. * At least 2 years of applied ML experience, including model development, evaluation, and dataset curation on complex scientific or biomedical data. * Strong proficiency with Python and the modern ML stack (e.g., PyTorch, scikit-learn) and SQL. * Experience turning complex, heterogeneous experimental data into robust features and training sets, with exposure to cloud training and data infrastructure. * Sound understanding of evaluation, validation, and the risks of leakage and distribution shift. * Ability to collaborate effectively with experimental scientists and modeling partners in a matrixed R&D environment. Preferred * Experience with biologics, antibody/protein sequence models, or protein language models. * Experience with active learning, Bayesian optimization, or sequence-based generative models for molecular design. * Familiarity with biophysical/assay data and developability endpoints. * Experience with MLOps, experiment tracking, and model monitoring. * Familiarity with how ontologies or knowledge graphs support data reuse and AI-ready datasets. ## Description Please note that this role is available across multiple countries and may be posted under different requisition numbers to comply with local requirements. While you are welcome to apply to any or all of the postings, we recommend focusing on the specific country(s) that align with your preferred location(s): USA - Requisition Number: R-095854 Spain - Requisition Number: R-096793 Why this role matters: Biologics Discovery is generating rich, fast-growing data across assays, sequences, and modalities, and the opportunity now is to make that data fully model-ready and seamlessly available for ML. This role ensures biologics data is structured for training, and that applied ML on discovery data helps scientists prioritize molecules, flag risks, and generate hypotheses earlier - strengthening the interface to ISD's models rather than duplicating them. Position Summary You will design robust featurization and dataset curation, build and evaluate applied models on biologics assay, biophysical, and sequence/construct data, and define evaluation frameworks that keep models trustworthy. You operate at the interface between our data-generating and data-infrastructure partners and In Silico Discovery (ISD), ensuring the datasets and features you create strengthen ISD's molecular property models. This is an opportunity to shape how AI learns from every biologics experiment., Featurization & Model-Ready Data * Develop featurization and model-ready datasets from antibody/protein sequence, construct, assay, and biophysical data. * Work with data engineers to specify the features, labels, and levels of aggregation that models need, preserving raw representations where information matters. * Curate, document, and version datasets so modeling is reproducible and traceable. Applied ML & Evaluation * Develop featurization and model-ready datasets from antibody/protein sequence, construct, assay, and biophysical data. * Work with data engineers to specify the features, labels, and levels of aggregation that models need, preserving raw representations where information matters. * Curate, document, and version datasets so modeling is reproducible and traceable. Partnership, Rigor & Growth * Collaborate with ISD to hand off standardized, traceable training datasets and align on where Data Science enables versus where ISD owns modeling. * Partner with Discovery scientists to frame ML problems around real decision points in the design-make-test-learn (DMTL) cycle. * Work closely with ontology and MLOps colleagues so datasets carry consistent semantics and models move reliably from development into use. * Champion reproducibility, documentation, and responsible AI. Why This Role Is Unique This is an opportunity to apply ML where it truly moves the needle in biologics discovery - grounded in real assay and sequence data, tightly partnered with world-class molecular modeling, and with real room to grow your scope, technical leadership, and impact as you build a track record of delivery., This position will be based at one of our office locations in either Spring House, PA (strongly preferred), Titusville, NJ, or Raritan, NJ, USA; or Madrid, Spain. (No remote option.) Johnson & Johnson is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, age, national origin, disability, protected veteran status or other characteristics protected by federal, state or local law. We actively seek qualified candidates who are protected veterans and individuals with disabilities as defined under VEVRAA and Section 503 of the Rehabilitation Act. If you are under 18 years of age, you (the candidate) may need to obtain the necessary working papers or other documentation required by state law to start the assignment, as well as get a parent's consent for the background check. Johnson & Johnson is committed to providing an interview process that is inclusive of our applicants' needs. If you are an individual with a disability and would like to request an accommodation, external applicants please contact us via , internal employees contact AskGS to be directed to your accommodation resource. The anticipated base pay range for this position is $109,000 to $174,800. The Company maintains highly competitive, performance-based compensation programs. Under current guidelines, this position is eligible for an annual performance bonus in accordance with the terms of the applicable plan. The annual performance bonus is a cash bonus intended to provide an incentive to achieve annual targeted results by rewarding for individual and the corporation's performance over a calendar/performance year. Bonuses are awarded at the Company's discretion on an individual basis. Employees and/or eligible dependents may be eligible to participate in the following Company sponsored employee benefit programs: medical, dental, vision, life insurance, short- and long-term disability, business accident insurance, and group legal insurance. Employees may be eligible to participate in the Company's consolidated retirement plan (pension) and savings plan (401(k)). Employees are eligible for the following time off benefits: Vacation - up to 120 hours per calendar year Sick time - up to 40 hours per calendar year Holiday pay, including Floating Holidays - up to 13 days per calendar year of Work, Personal and Family Time - up to 40 hours per calendar year Additional information can be found through the link below. https://www.careers.jnj.com/employee-benefits The compensation and benefits information set forth in this posting applies to candidates hired in the United States. Candidates hired outside the United States will be eligible for compensation and benefits in accordance with their local market. #LI-SL #JNJDataScience #JNJIMRND-DS #JRDDS #LI-Hyrbid # 3, Job Title: Data Scientist (Locations available in: Warminster, PA, Hilliard, OH, Plymouth, MI, and Burnsville, MN) Department: Finance Accountability: This Position Reports to … + 1 month ago + ## Related Videos - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Overview of Machine Learning in Python](https://www.wearedevelopers.com/videos/840-overview-of-machine-learning-in-python) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)