Principal Engineer - Data Intelligence

Yahoo
Nashville, TN, United States
14 days ago
Apply on www.jofdav.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$143,625.0 - $299,375.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Big Data BigQuery Distributed Data Store Distributed Systems Data Flow Control Graph Database Data Intelligence Mobile Application Software Python (Programming Language) Pair Programming Search Technologies
+8 more
Software Engineering Multi-Agent Systems Apache Spark Generative AI Apache Kafka Build Tools Virtual Agents Stream Processing

Job description

It takes powerful technology to connect our brands and partners with an audience of hundreds of millions of people. Whether you’re looking to write mobile app code, engineer the servers behind our massive ad tech stacks, or develop algorithms to help us process trillions of data points a day, what you do here will have a huge impact on our business-and the world., We are seeking a Principal Engineer (IC5) to lead the architecture and development of AI-powered data intelligence platforms operating at Yahoo scale.

This role sits at the intersection of large-scale distributed data systems, enterprise knowledge, Retrieval-Augmented Generation (RAG), and agentic AI. You will define how AI reasons over Yahoo’s data ecosystem to help teams detect anomalies, diagnose incidents, understand dependencies, discover trusted data, optimize infrastructure, and automate complex operational workflows.

As a Principal Engineer, you will establish technical direction and reference architectures for enterprise AI systems while building reusable platform capabilities that can be adopted across engineering, analytics, governance, and business organizations.

You will also help shape how Yahoo engineers use AI to build software-leveraging AI-assisted development, evaluation, and automation to accelerate engineering while maintaining rigorous standards for correctness, reliability, security, and operational safety.

What you’ll build

Systems that answer questions like:

  • Why did engagement decline after last week’s deployment?
  • What breaks downstream if we change this schema?
  • What’s the root cause of this data quality incident?
  • Which datasets should I trust for this decision?
  • Where can we cut capacity waste without slowing consumer queries?

Delivered through natural language interfaces over enterprise metadata and operational telemetry, serving executives, data producers, data consumers, and governance teams.

Responsibilities

  • Define technical vision and architecture for Yahoo’s AI-powered data intelligence platform.
  • Architect petabyte-scale batch and streaming systems with Beam, Dataflow, Spark, and Kafka.
  • Build production RAG and multi-agent systems: semantic metadata integration, agent orchestration.
  • Establish evaluation frameworks - retrieval quality, agent accuracy, hallucination rate, operational safety - with automated benchmarking, so we ship AI we can trust.
  • Use AI development tools to accelerate design, implement the bar for how the org validates AI-generated code and agent plans.
  • Drive adoption of RAG and agentic patterns across engineers.
  • Partner with product, governance, infrastructure, and exit-form priorities., * Agentic AI frameworks (LangGraph, CrewAI, AutoGen, Semant)
  • GCP: Vertex AI, Gemini, BigQuery, Vector Search, Dataflow
  • Enterprise metadata, lineage, governance, or observability
  • Knowledge graphs and semantic reasoning systems.
  • Daily use of AI pair-programming tools (Claude, Copilot, g workflows).
  • Led large technical initiatives influencing strategy across multiple teams.

The material job duties and responsibilities of this role include those listed above as well as adhering to Yahoo policies ; exercising sound judgment ; working effectively, safely and inclusively with others ; exhibiting trustworthiness and meeting expectations ; and safeguarding business operations and brand integrity.

At Yahoo, we offer flexible hybrid work options that our employees love! While most roles don’t require regular office attendance, you may occasionally be asked to attend in-person events or team sessions. You’ll always get notice to make arrangements. Your recruiter will let you know if a specific job requires regular attendance at a Yahoo office or facility. If you have any questions about how this applies to the role, just ask the recruiter!

Requirements

  • 10+ years building distributed systems and large-scale data platforms, or equivalent experience.
  • Degree in CS, AI/ML, Engineering, Math, or related field,
  • Deep Python expertise and strong software engineering fun
  • Production experience with batch and streaming processing (Kafka, or equivalent).
  • Shipped and operated RAG applications in production; search, and reranking.
  • Track record evaluating and validating AI-generated code correctness and operational safety.
  • Strong communication and technical leadership across cross-functional stakeholders.

Benefits & conditions

The compensation for this position ranges from $143,625.00 - $299,375.00/yr and will vary depending on factors such as your location, skills and experience.The compensation package may also include incentive compensation opportunities in the form of discretionary annual bonus or commissions. Our comprehensive benefits include healthcare, a great 401k, backup childcare, education stipends and much (much) more.

About the company

The Consumer Data’s Foundations & AI Platform Engineering team builds the foundational intelligence layer powering Yahoo’s global ecosystem, processing over 100 PB of data daily for more than a billion users. We combine large-scale distributed data systems, enterprise knowledge and metadata, and modern AI to transform how data is discovered, understood, trusted, governed, and operated across the company.

Our mission is to build the intelligence and reasoning layer over Yahoo’s enterprise data ecosystem-connecting metadata, lineage, operational telemetry, business context, and governance signals so teams can understand what is happening across complex data systems and take action with confidence.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy ¡ LIVE

2:30 min

Leveraging BigQuery ML for scalable SQL-based segmentation experiments

Julian Joseph ¡ LIVE

48 sec

Exploring alternative build tools and experimental web components

Sasha Shynkevich ¡ LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou ¡ Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy ¡ LIVE

2:09 min

Configuring IDEs and build tools for Java 17

Daniel Strmečki · LIVE

Videos

See all

Related articles

See all