> Markdown version of [/jobs/ext/2035950-lead-data-scientist-ai-labs](https://www.wearedevelopers.com/jobs/ext/2035950-lead-data-scientist-ai-labs). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Scientist, AI Labs - **Company:** NPR - **Location:** Washington, DC, United States - **Experience:** Expert - **Salary:** $164,000.0 - $201,000.0 - **Contract:** Permanent contract - **Skills:** Sql Data Warehouse, Artificial Intelligence, Amazon Web Services, Microsoft Azure, Big Data, BigQuery, Continuous Integration, Database Search Engine, Python (Programming Language), Machine Learning, Natural Language Processing, Open Source Technology, Recommender Systems, Regression Testing, Search Technologies, SQL Databases, Speech Recognition, Delivery Pipeline, Large Language Models, Snowflake, Prompt Engineering, Model Validation, Information Technology, Data Management, Machine Learning Operations, Api Design, Data Pipelines - **Published:** August 12, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88008034/1 ## About the Role * 8+ years of professional experience in Data Science, Machine Learning, or Natural Language Processing (NLP) shipping production-grade systems. * Proven track record of applying, fine-tuning, and evaluating Large Language Models (LLMs) and foundational architectures. * Experience with Automatic Speech Recognition (ASR), diarization, and audio preprocessing pipelines. * Demonstrated experience designing and maintaining large-scale data architectures, vector databases, and semantic search pipelines. * Hands-on experience designing, deploying, and optimizing production-grade recommendation engines or personalization systems at scale. * Practical, hands-on experience building machine learning workflows within major cloud environments (Azure, AWS or GCP). * Experience successfully navigating matrixed, cross-disciplinary collaboration between technical engineering teams and non-technical stakeholders. * Ability to identify common AI pitfalls, such as hallucinations or formatting errors, and design robust algorithmic workarounds., * Experience working with large volumes of unstructured multimedia, digital audio processing, or automated speech-to-text workflows. * Prior experience working within a media organization, digital newsroom, or public service institution. * Experience integrating automated LLM evaluation gates and regression testing directly into CI/CD deployment pipelines, * Deep technical expertise in Python, ML and agentic engineering frameworks, vector search engines and two-stage retrieval architectures. * Strong proficiency in SQL and cloud data warehouse ecosystems (such as BigQuery or Snowflake). * Deep understanding of Retrieval-Augmented Generation (RAG) patterns and semantic search tools. * Professional-level familiarity with model evaluation methodologies, prompt engineering techniques, and API integration workflows. * Strong commitment to algorithmic ethics, data privacy compliance, and methods for identifying or mitigating bias. Education Requirements * Advanced degree (PhD preferred) in Computer Science, Data Science or Machine Learning * Advanced academic research in relevant fields ## Description Across our organization, we're building a workplace where collaboration is essential, diverse voices are heard, and inclusion is the key to our success. We are committed to doing the right thing in our journalism and in every role at NPR.This means that integrity, adherence to our ethical standards, and compliance with legal obligations are fundamental responsibilities for every employee at NPR. Intro to Position As the Lead Data Scientist for the AI Labs team, you will serve as the technical and ethical anchor for NPR's artificial intelligence initiatives. You will lead data science expertise for a content metadata overhaul to power new audience-focused personalization engines. Rather than building models from scratch, you will focus on fine-tuning, evaluating, and scaling existing foundational models, as well as deploying machine learning algorithms for public media.. You will collaborate extensively across the organization to align architectures, ensure secure cloud deployments, and protect intellectual property. This position demands a high focus on accuracy, journalistic ethics, data privacy, and responsible scaling. Responsibilities * Lead the selection, fine-tuning, and optimization of open-source and proprietary LLMs tailored to NPR's unique content voice. * Design and architect automated machine learning pipelines to transform decades of unstructured audio, transcripts, and text to support automated semantic metadata generation. * Collaborate with the product, design and engineering teammates to translate user needs into production-ready data science workflows. * Architect recommendation frameworks that leverage enriched metadata to drive deep, style-based audience personalization while preserving editorial curation. * Partner with Data Products to construct clean, self-service data pipelines and audience analytics models inside BigQuery. * Support newsroom research through the prototyping, validation, and development of API-driven tooling and secure database search models. * Establish strict evaluation, testing, and benchmarking frameworks to guarantee model outputs meet NPR's standards for factual accuracy and neutrality. * Proactively identify, audit, and mitigate algorithmic bias in metadata generation and audience discovery systems. * Ensure all AI applications scale securely and cost-effectively, balancing computational efficiency with rigorous data privacy guardrails. * Collaborate with growth and data platforms to leverage content metadata for user lifecycle retention and smart audience segmentation. ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Making Data Warehouses fast. A developer's story.](https://www.wearedevelopers.com/videos/302-making-data-warehouses-fast-a-developer-s-story) - [Data Governance in the Era of AI](https://www.wearedevelopers.com/videos/1622-data-governance-in-the-era-of-ai) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career)