Pre-Training Text Data

Microsoft
Redmond, United States of America
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Intermediate
Compensation
$ 235K

Job location

Irving, United States of America

Tech stack

Artificial Intelligence
Data analysis
Big Data
Information Engineering
Data Transformation
Python
Microsoft Office
NumPy
Software Engineering
Management of Software Versions
Data Ingestion
Spark
Pandas
Information Technology
Integration Frameworks
Data Pipelines
Apache Beam

Job description

We are seekingengineers and researchersto join our Pretraining Text Data team, where we are building the next generation of foundation large language models. If you are passionate about designing and curating high-quality datasets to power frontier AI models, this role is for you.

In this role,you'llwork at the intersection of data and innovation-collaborating with scientists, engineers, and annotators to curate, analyze, and evaluate diversetextdatasetscritical to model development. You will lead efforts to:

  • Develop novel data collection strategies

  • Improve dataset quality and integrity

  • Understand data-driven model behaviors

  • Train models to understand the impact of data and data mixes

  • Align datasets with ethical and societal values

This is a cross-disciplinary, high-impact role ideal for engineersand researcherswho want to push the boundaries of what AI can learn from data.

Microsoft Superintelligence Team

MicrosoftSuperintelligence team's mission is to empower every person and every organization on the planet to achieve more. Asemployeeswe come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

This role is part of Microsoft AI's Superintelligence Team. The MAIST is a startup-like team inside Microsoft AI, created to push the boundaries of AI toward Humanist Superintelligence-ultra-capable systems thatremaincontrollable, safety-aligned, and anchored to human values. Our mission is to create AI that amplifies human potential while ensuring humanityremainsfirmly in control. We aim to deliver breakthroughs thatbenefitsociety-advancing science, education, and global well-being.

We'realso fortunate to partner with incredible productteamsgiving our models the chance to reach billions of users and createimmensepositive impact. Ifyou'rea brilliant, highly-ambitiousand low ego individual,you'llfit right in-come and join us as we work on our next generation of models!

Starting January 26, 2026, MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary byjurisdiction.

Responsibilities

  • Create high-quality datasets for training and evaluation; run experiments on new datasets (data ablations) to assess their impact anddeterminethe most effective data.
  • Develop and maintain scalable data pipelines for text data ingestion, preprocessing, filtering, and annotation.
  • Analyze real-world text datasets to assess quality, diversity, relevance, and identify areas for improvement.
  • Build lightweight tools and workflows for dataset auditing, visualization, and versioning.
  • Collaborate with Safety, Ethics, and Governance teams to ensure datasets meet standards for quality, privacy, and responsible AI practices.
  • Embody our cultureandvalues.

Requirements

  • Bachelor's Degree in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or related technical discipline AND 4+ years technical engineering experience with coding in languages including, but not limited to, Python and common data libraries (Pandas, NumPy, etc.)
  • OR equivalent experience., * Master's Degree in in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or related technical discipline AND 8+ years technical engineering experience with coding in languages including, but not limited to, Python and common data libraries (Pandas, NumPy, etc.)
  • OR Bachelor's Degree in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or related technical discipline AND 12+ years technical engineering experience with coding in languages including, but not limited to, Python and common data libraries (Pandas, NumPy, etc.)
  • OR equivalent experience.
  • 2+ years of experience in data analysis or data engineering, including work with large-scale datasets that are unstructured or semi-structured.
  • Proficiency in statistics and exploratory data analysis methods.
  • Familiarity with data processing frameworks such as Spark, Ray, or Apache Beam.
  • Ability to communicate technical findings clearly to research and product teams.

Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800.00 - $234,700.00 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200.00 - $261,000.00 per year.

About the company

Microsoft is a global technology company headquartered in Redmond, Washington. Our mission is to empower every person and every organization on the planet to achieve more. We develop, license, and support a wide range of software products, services, and devices that help individuals and businesses realize their full potential.

Our flagship products include the Microsoft 365 productivity cloud, Windows operating system, Azure cloud platform, and Dynamics 365 business applications. We are also a leader in areas such as artificial intelligence, cybersecurity, developer tools, and gaming through Xbox and Game Pass.

With operations in more than 190 countries and over 220,000 employees worldwide, Microsoft is committed to responsible innovation, inclusive economic growth, and sustainability. We work closely with governments, industries, and communities to ensure that technology serves the public good and helps address some of the world’s most pressing challenges.

As we celebrate our 50th anniversary in 2025, we continue to look forward—investing in AI, cloud, and quantum computing to shape the future of work, education, and society at large scale.

Apply for this position