> Markdown version of [/jobs/ext/1902759-principal-data-scientist-immunology](https://www.wearedevelopers.com/jobs/ext/1902759-principal-data-scientist-immunology). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Data Scientist - Immunology - **Company:** Johnson & Johnson - **Location:** San Diego, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $117,000.0 - $201,250.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Analysis, Microsoft Azure, Bioinformatics, Health Informatics, Clinical Data Repository, Computer Programming, Computer Literacy, Continuous Delivery, Continuous Integration, Data Dictionary, Data Files, Data Governance, Data Visualization, DevOps, Document Management Systems, EHealth, Graph Database, Health Information Technology, Interoperability, Linked Data, Natural Language Processing, Neo4j, Resource Description Framework (RDF), Requirements Management, Semantic Web, SPARQL, SQL Databases, Data Processing, Data Storage Technologies, System Availability, Gitlab, Git, Containerization, Information Technology, Data Lineage, Graphql, Restful APIs, Docker, Jenkins - **Published:** July 31, 2026 - **Apply:** https://www.careerbuilder.com/job-details/principal-data-scientist-immunology-2-positions-san-diego-ca--6f1af080-da1f-43cc-9bc2-fbce56a956ae ## About the Role * Desired Ph.D. or master's degree in bioengineering, computer science, IT, bioinformatics, physics, mathematics, or related fields, emphasis on semantic technologies for biomedical application. * 5+ years professional experience in health informatics. * Demonstrated experience in large-scale knowledge graphs construction, ontology development, pharmaceutical or healthcare domains integration. * Programming background in parser combinators, natural language processing, and linked data (RDF Triple Stores and property graphs). * Proficiency in semantic web technologies (e.g. SPARQL, RDF, OWL), familiarity with graph databases (Neo4j, Amazon Neptune). * Proven work with complex biomedical datasets (e.g. clinical, genomics, proteomics) * Proficiency in various data storage solutions (SQL, key-value, column, document, graph stores) and data modeling techniques (semantic data, ontologies, taxonomies). * Experience in CI/CD implementations, git usage, CI/CD stacks (Jenkins, GitLab, Azure DevOps), DevOps tools, metrics/monitoring, and containerization technologies (Docker, Singularity). * Demonstrated stakeholder management capabilities- including requirements gathering, business analysis and planning. Must have the capacity to translate discussions into user requirements and project plans. * Ability to manage a numerous projects simultaneously, prioritize work, exhibit organizational skills and flexibility to deliver maximum business value. * Willingness to conduct periodic travel ( This position will be located on site at one of our campuses in either Spring House PA, Cambridge MA, Titusville NJ, Raritan NJ, or San Diego, CA (NO fully remote option available). Occasional travel for cross-functional workshops, design sessions, and team meetings may be required., Preferred Skills: Advanced Analytics, Coaching, Critical Thinking, Data Analysis, Data Privacy Standards, Data Quality, Data Reporting, Data Savvy, Data Science, Data Visualization, Digital Fluency, Econometric Models, Organizing, Process Improvements, Strategic Thinking, Technical Credibility, Workflow Analysis Remote Skills: Artificial Intelligence (AI), Bioengineering, Bioinformatics, Biology, Biomedical Software, Biomedicine, Borland ObjectWindows Library (OWL) Programming Libraries, Business Analysis, Business Plan, Clinical Data, Clinical Research, Coaching, Compensation and Benefits, Computer Science, Construction, Continuous Deployment/Delivery, Continuous Integration, Cross-Functional, Data Analysis, Data Modeling, Data Processing, Data Quality, Data Science, Data Sets, Data Storage, Data Visualization, Database Administration, DevOps, Disease, Disease Prevention and Control, Diversity, Docker, Document Management, Drug Development, Econometric Modeling, Establish Priorities, Genomics, Git, Graph Database Data Format, GraphQL, Health Information Technology, Health Science, Healthcare, High Availability, Immunology, Internet Technology, Interoperability, Jenkins, Mathematics, Medicine, Metrics, Microsoft Windows Azure, Multitasking, Natural Language Processing (NLP), Neo4j, Ontology, Organizational Skills, Physics, Predictive Modeling, Process Improvement, Product Lifecycle, Project Planning, Proteomics, RDF (Resource Description Framework), REST (Representational State Transfer), Requirements Management, Research & Development (R&D), SPARQL, SQL (Structured Query Language), Taxonomies, Technical Strategy, Technical/Engineering Design, Willing to Travel, Workflow Analysis ## Description * Be a key contributor to the design and implementation of a scalable knowledge graph infrastructure focused on data standardization and interoperability, focusing on Immunology R&D data. * Apply graph-based data modeling for efficient Immunology R&D organization, integration and retrieval to ensure system flexibility and long-term maintainability. * Work with a larger community of Data Scientists, Clinical Scientists, and Discovery Scientists to standardize, curate and create AI-Ready data sets. * Curate and extend ontologies for clear mapping into established biomedical ontologies and controlled terminologies using resource description framework (RDF) standards. * Work with SPARQL/GraphQL/REST services; develop ingestion and curation pipelines to ingest, normalize and map concepts across data sources. * Extend and curate Immunology R&D-relevant ontologies (e.g., diseases, drugs, targets, pathways, etc.) and maintain synonyms, cross-references, and provenance. * Partner with cross-functional teams to enable NLP/RAG over graphs, features for predictive modeling and terminology services for search and study design tools. * Work with Data Science & Digital Health colleagues, IT and DevOps teams to deploy and manage the graph database infrastructure, focusing on high availability, scalability, and recovery operations specifically geared toward Immunology R&D needs and applications. * Draft and manage documentation, such as data dictionaries, data lineage, and data flow diagrams, to facilitate understanding of the knowledge graph. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Best Coding Boot Camps in Germany](https://www.wearedevelopers.com/magazine/237-best-coding-boot-camps-in-germany) - [The Fastest-Growing Tech Sectors to Look Out for in 2025](https://www.wearedevelopers.com/magazine/373-the-fastest-growing-tech-sectors-to-look-out-for-in-2025) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe)