> Markdown version of [/jobs/ext/1945812-principal-scientist-oncology-data-science-translational-science](https://www.wearedevelopers.com/jobs/ext/1945812-principal-scientist-oncology-data-science-translational-science). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Scientist, Oncology Data Science (Translational Science) - **Company:** Glaxosmithkline PLC - **Location:** Durham, NC, United States (Remote available) - **Salary:** $121,275.0 - $202,125.0 - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Artificial Intelligence, Data Analysis, Automation of Tests, Clinical Data Repository, Code Reuse, Computational Biology, Continuous Integration, Information Engineering, Python (Programming Language), Machine Learning, Modular Design, Software Engineering, Data Processing, Pytorch, Deep Learning, Information Technology, Stable Diffusion, Virtual Agents, Software Version Control, GXP - **Published:** August 6, 2026 - **Apply:** https://gsk.wd5.myworkdayjobs.com/GSKCareers/job/USA---Pennsylvania---Upper-Providence/Principal-Scientist--Oncology-Data-Science--Translational-Science-_445634 ## About the Role * PhD (or equivalent experience) in a quantitative field (Applied ML, Computer Science, Physics, Systems/Computational Biology, or equivalent) with 1+ years of industry or productive post-doctoral academic experience. * Experience with deeply embedded in cancer / computational biology, with a strong understanding of tumor microenvironment dynamics and high dimensional datasets. * Experience with analytical and modelling skills, including expertise in statistical and machine learning approaches. * Experience with the analysis of single cell omics data. * Experience in one or more of the following: statistical modelling of functional genomics screening datasets (e.g., bulk CRISPR screens, Perturb-seq) or spatial omics. * Experience in Python and deep learning frameworks (PyTorch) for data processing and machine learning model development, with a strong grasp of software engineering fundamentals (e.g., version control, modular design, CI/CD). Preferred Qualification If you have the following characteristics, it would be a plus: * Experience with multi-modal integration, including spatial transcriptomics/proteomics, histopathology, and single cell omics data. * Experience working with longitudinal clinical health record trajectory data. * Familiarity with R for specialized statistical modelling. * Experience with causal inference and individual treatment effect modelling. * Experience with AI agent-driven workflows and coding tools. * Experience with generative deep learning approaches, including flow matching, diffusion and causal transformer models. * Excellent written and oral communication skills, with a proven ability to present complex computational concepts to technical and non-technical stakeholders. Work model: This role is hybrid. You will balance on-site collaboration with focused remote work., Applied Statistics, Data Analysis, Data Engineering, Data Science, Datasets, Drug Development, Drug Discovery Process, Drug Target Identification, Genetic Analysis, Genomic Analysis, Machine Learning (ML), Software Engineering ## Description The GSK Oncology Data Science team in R&D Translational Science is seeking a Translational AI scientist to build ML applications for a long-sought-after problem: if we alter a patient tumor's molecular state in silico, can we predict how their clinical trajectory will change? To tackle this problem, you will integrate and validate multimodal foundation models; bridge functional genomics, spatial omics, and real-world data; and apply cutting edge causal inference techniques. We operate with high velocity at the intersection of machine learning, causal inference, functional genomics, spatial biology, and real-world clinical data; your expertise, execution, technical leadership and communication will drive our efforts to bring the right therapies to the right patients. Responsibilities This role will provide YOU the opportunity to lead key activities to progress YOUR career. These responsibilities include some of the following: * Own the pipeline and develop advanced ML architectures to integrate complex multimodal datasets, including single-cell, spatial omics, histopathology, functional genomics, and real-world clinical data. * Partner closely with wet-lab scientists, clinicians, and pathologists to validate machine learning models, including in-silico perturbations within the tumor microenvironment against ground-truth data (counterfactual validation). * Develop approaches to extract interpretable features from models to generate testable oncological hypotheses and link insights to clinical pipeline decisions such as asset prioritization and patient subpopulation selection. * Contribute clean, reproducible tooling to cross-team frameworks. We enforce good engineering practices in our research-utilizing code architecture planning, clean code and automated testing to build trustworthy, reusable code. * Maintain cutting edge knowledge of advancements, share with the team and maintain our team as a thought leader through publications in high-impact venues and engaging with the broader community. ## Related Videos - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Introduction to sCrypt - a smart contract language for Bitcoin SV](https://www.wearedevelopers.com/videos/24-introduction-to-scrypt-a-smart-contract-language-for-bitcoin-sv) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [The Fastest-Growing Tech Sectors to Look Out for in 2025](https://www.wearedevelopers.com/magazine/373-the-fastest-growing-tech-sectors-to-look-out-for-in-2025) - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies) - [The Prompt Engineer ✍️](https://www.wearedevelopers.com/magazine/216-the-prompt-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)