Data Flywheel Infrastructure Engineer

Microsoft
San Francisco, CA, United States
8 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$119,800.0 - $234,700.0
Working hours
Regular working hours
Job source

Tech stack

Training Data Artificial Intelligence Computer Engineering Data Deduplication Information Engineering Data Files Data Governance Data Infrastructure Data Mining Data Security Information Lifecycle Management Python (Programming Language)
+11 more
Microsoft Office Software Engineering SQL Databases Management of Software Versions Data Ingestion Large Language Models Apache Spark Information Technology Apache Flink Data Pipelines Data Generation

Job description

We are looking for a Data Flywheel Infrastructure Engineer to build the infrastructure that continuously turns 1P data, 3P data, model signals, evaluation results, and synthetic data into high-quality training data for frontier LLM and multimodal models. This role owns the systems connecting: Data Acquisition Governance & Compliance Curation Training Evaluation Failure Mining Data Improvement A critical part of the role is enabling aggressive data iteration while ensuring that every dataset is secure, policy-compliant, rights-aware, traceable, and auditable.

Starting January 26, 2026, MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction.

This role is part of Microsoft AI’s Superintelligence Team. The MAIST is a startup-like team inside Microsoft AI, created to push the boundaries of AI toward Humanist Superintelligence-ultra-capable systems that remain controllable, safety-aligned, and anchored to human values. Our mission is to create AI that amplifies human potential while ensuring humanity remains firmly in control. We aim to deliver breakthroughs that benefit society-advancing science, education, and global well-being.

We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in-come and join us as we work on our next generation of models!

Responsibilities

  1. Build 1P & 3P Data Flywheel Infrastructure Build scalable systems for ingesting, processing, curating, versioning, and serving first-party and third-party data for pre-training and post-training. Connect model failures, evaluations, and product signals back into targeted data acquisition, generation, and improvement workflows.

  2. Own Data Governance, Security & Compliance Infrastructure Build governance and policy enforcement directly into the data platform, including:

*

  • Data provenance and lineage
  • Usage rights, licensing, and consent metadata
  • PII / sensitive-data detection and protection
  • Access control and data isolation
  • Retention and deletion enforcement
  • Geographic and regulatory restrictions
  • Dataset approval and audit workflows
  • Training eligibility and purpose-based usage controls
  1. Build Policy-Aware Data Acquisition & Curation Systems Develop automated pipelines for 1P and 3P data ingestion, classification, filtering, deduplication, quality scoring, semantic enrichment, and dataset construction. Make governance policies machine-enforceable so that data can automatically be included, excluded, quarantined, or restricted based on its origin, license, sensitivity, consent, geography, and intended model use.

  2. Build Evaluation-to-Data Feedback Loops Convert model evaluations and real-world failure signals into actionable data tasks through failure clustering, hard-example mining, long-tail discovery, capability-gap detection, and targeted dataset generation. Enable rapid iteration from: Model Failure Data Gap Data Intervention Training Evaluation

  3. Build Synthetic & AI-Native Data Pipelines Use LLMs, VLMs, and Agents to automate data generation, labeling, filtering, quality validation, enrichment, and transformation. Maintain clear provenance between human-created, first-party, third-party, model-generated, and derived data, and enforce appropriate policies across each category.

  4. Build Data Quality, Attribution & Observability Develop metrics and infrastructure to measure dataset quality, coverage, diversity, contamination, duplication, policy compliance, and contribution to model capability improvements. Enable researchers to understand which data improves which capabilities and under what governance constraints.

Requirements

Required Master’s Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 3+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor’s Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience. Software Engineering experience using Python,SQL, Spark/Flink/Ray

Preferred

  • Experience building AI training-data governance platforms, including provenance, licensing/rights metadata, consent management, PII handling, policy enforcement, or auditable lineage.
  • Experience managing third-party datasets, data partnerships, licensed content, or externally sourced data with complex contractual and usage restrictions.
  • Experience building privacy- and security-aware systems for first-party product or user data, including isolation, access controls, retention/deletion, and purpose limitation.
  • Experience with data clean rooms, privacy-preserving processing, de-identification, confidential computing, or secure data collaboration.
  • Experience building evaluation failure mining data generation training feedback loops.
  • Experience with synthetic data, model graders, reward signals, hard-example mining, active learning, or data-mixture optimization.
  • Experience with multimodal or agentic datasets including text, image, video, audio, web, GUI, tool-use, or interaction trajectories.
  • Understanding of Modern LLM training workflows including Pre-training, SFT, RL/post-training, evaluation, and synthetic data.
  • Strong understanding of data governance, security, privacy, provenance, access control, and data lifecycle management.

Data Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:55 min

Infrastructure challenges when combining Kafka with Apache Flink

Bobur Umurzokov · LIVE

2:31 min

Data mining literary works for language patterns

Jen Looper Jen Looper · Europe 2026 Virtual

3:38 min

Reusing email software standards for HTTP file uploads

Imran Nazar · World Congress 2023

1:59 min

Evolving roles in AI driven software teams

Ignacio Riesgo Ignacio Riesgo +1 · World Congress 2024

1:09 min

Evaluating mature stream processing frameworks for production systems

Soroosh Khodami Soroosh Khodami · World Congress 2024

Videos

See all

Related articles

See all