machine learning engineer qa artificial intelligence architecture

Airbnb
United States
7 days ago
Apply on www.workingnomads.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$244,000.0 - $305,000.0
Working hours
Regular working hours

Tech stack

A/B Testing Agile Methodology Artificial Intelligence Automation of Tests Data Validation Data Infrastructure E-Business Machine Learning Regression Testing Azure Machine Learning Management of Software Versions Large Language Models
+4 more
Generative AI Information Technology Power Analysis (Cryptography) Data Pipelines

Job description

AI and ML are at the heart of the Airbnb product. From Trust to Payments, and from Customer Service to Marketing, we rely on ML to ensure that guests and hosts have the best possible experience with Airbnb.

The Core ML team is responsible for driving CSxAI (Customer Support x Artificial Intelligence) initiatives by adopting Generative AI technologies to enable an intelligent, scalable, and exceptional service experience. The team develops and enhances AI models, ML services, and tools including LLM fine-tuning and optimization, RAG/Search, LLM evaluation and testing automation, feedback-based learning, and guardrails for a wide range of applications at Airbnb.

The richness of Airbnb’s data, the complexity of its marketplace, and the variety innate in our product mean that we need to operate at the state of the art of AI practice. We are committed to long-term innovation to solve complex problems, and to do that we need experienced ML

The Difference You Will Make:

In this Senior Staff role, you will set technical direction and lead execution for ML evaluation and the end-to-end data flywheel powering CSxAI products (e.g., assistive agents, issue resolution, and tooling). Your work will define how we measure quality, how we turn feedback into learning signals, and how we continuously improve models and products safely and efficiently. You will partner closely with product, engineering, design, operations to build evaluation systems that are trusted, scalable, and actionable - connecting offline metrics to online outcomes.

A Typical Day:

  • Define evaluation strategy and success metrics for GenAI systems, aligning offline evaluation with online business and customer experience outcomes.
  • Build and scale evaluation frameworks (golden sets, synthetic data, automated regressions, rubric-based grading, LLM-as-judge where appropriate) with strong controls for bias, drift, and reliability.
  • Design the data flywheel: instrumentation, feedback collection, data quality checks, labeling strategy, dataset versioning, and governance to support continuous improvement.
  • Lead cross-functional quality initiatives across product, ops, and engineering, driving clarity on what “good” looks like and how teams act on evaluation results.
  • Develop and productionize pipelines for dataset creation, model monitoring, evaluation-at-scale, and continuous testing (pre-deploy and post-deploy).
  • Drive technical decisions and architecture for evaluation and data infrastructure, balancing speed, rigor, cost, and safety.

Requirements

  • Educational Background: PhD in Computer Science, Mathematics, Statistics, or related technical field (or equivalent practical experience).
  • Industry Experience: 10+ years building, testing, and shipping ML/AI systems end-to-end; including 2+ years of experience with GenAI/LLM systems in production.
  • Leadership Experience: 5+ years leading large, ambiguous technical initiatives as a senior IC, influencing roadmap and engineering/science direction across teams.
  • Technical Proficiency:
  • Deep expertise in evaluation methodology (offline/online alignment, metric design, human-in-the-loop evaluation, A/B testing, power analysis, regression testing).
  • Hands-on experience with GenAI systems, including orchestration, retrieval, tool calling, memory, etc.
  • Experience building data pipelines and quality systems (labeling workflows, dataset curation, versioning, monitoring, and governance).
  • Solid ML fundamentals and best practices (model selection, training/serving, monitoring, reliability, and model lifecycle management).

Preferred Qualifications:

  • Customer Support Systems: Experience applying ML/AI to customer support workflows (e.g., agent assist, classification/routing, resolution recommendation, QA).
  • Infrastructure & Quality at Scale: Experience building robust evaluation platforms for agent behavior validation, safety/guardrails, and continuous improvement.
  • Agile Practice for Applied AI: Proven ability to take evaluation and data flywheel work from incubation to production, iterating quickly while maintaining scientific rigor.

Benefits & conditions

Our job titles may span more than one career level. The actual base pay is dependent upon many factors, such as: training, transferable skills, work experience, business needs and market demands. The base pay range is subject to change and may be modified in the future. This role may also be eligible for bonus, equity, benefits, and Employee Travel Credits. Pay Range $244,000-$305,000 USD Go ad-free with Premium ×, Our job titles may span more than one career level. The actual base pay is dependent upon many factors, such as: training, transferable skills, work experience, business needs and market demands. The base pay range is subject to change and may be modified in the future. This role may also be eligible for bonus, equity, benefits, and Employee Travel Credits. Pay Range $244,000-$305,000 USD

About the company

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.workingnomads.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:06 min

Elevating the QA engineering role for complex challenges

Ondřej Gróf Ondřej Gróf · World Congress 2026 Europe

54 sec

Generating multiple hook options for outreach A/B testing

Leandro Gomes da Silva Leandro Gomes da Silva · World Congress 2025

2:37 min

Tracing the evolution from early AI to generative AI

Mike Mike · World Congress 2025

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

56 sec

Performing local A/B testing across multiple AI agents

Julia Kasper · Coffee With Developers

2:14 min

Introduction to generative AI and content warnings

Cheuk Ho · World Congress 2023

Videos

See all

Related articles

See all