Machine Learning Engineer

OpenAI Inc.
Bellevue, WA, United States
4 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Automated Storage and Retrieval Systems Python (Programming Language) Machine Learning Large Language Models Generative AI Backend Information Technology Machine Learning Operations GPT

Job description

We are looking for a Machine Learning Engineer to lead the technical direction for ML-powered experimentation and insights capabilities. You will build production systems that learn from privacy-protected product and experimentation data to generate evidence-backed insights and support decision-making, and help teams decide which ideas are worth testing live.

This is an end-to-end, 0-to-1 role. You will work across ML modeling, retrieval and LLM systems, statistical methods, simulation, data and training pipelines, backend services, and user- and agent-facing product experiences. The hard part is not merely producing a plausible answer. It is making each insight and prediction traceable, calibrated, useful, and safe enough to influence real product decisions.

Live experiments remain the source of causal validation. You will design systems that make uncertainty explicit, backtest against historical outcomes, compare predictions with online results, learn from misses, and abstain when the evidence is weak. You will preserve clear review, permission, and approval boundaries as automation becomes more powerful.

You will collaborate closely with teams building ChatGPT, Codex, model measurement workflows, consumer products, Growth, business subscription experiences, developer products, and shared infrastructure. You will turn their most important learning and decision problems into general platform capabilities that can support the full company., * Set and execute the technical roadmap for Generative Insights and Predictive Experimentation, from early prototypes through production adoption.

  • Build cross-experiment learning systems that retrieve and synthesize historical experiments, detect recurring effects and segment behavior, reanalyze prior results when data or methods improve, and generate hypotheses with clear evidence and provenance.
  • Develop predictive models and simulation workflows, including simulation-based evaluation approaches, to estimate likely impact, affected segments, regression risk, and uncertainty before a full live experiment.
  • Create high-quality datasets and feature or retrieval pipelines from exposures, events, metrics, experiment metadata, and replay data, with strong lineage, freshness, privacy, and data-quality controls.
  • Establish rigorous evaluation through offline benchmarks, backtests, calibration, drift monitoring, prediction-to-outcome comparisons, and explicit failure or abstention behavior.
  • Turn models into durable product, API, and agent workflows that move from an insight to experiment design, approval-gated action, and measured learning.
  • Partner deeply with data science and product teams on experiment design, causal inference, sequential decision-making, variance reduction, and the boundary between prediction and causal evidence.
  • Build reliable services and intuitive workflows so sophisticated ML capabilities are understandable and useful to teams making high-stakes product decisions.
  • Provide technical leadership across engineering, product, data science, and research partners, and raise the bar for production ML quality across the platform.

Requirements

  • Have led ambiguous 0-to-1 production ML products where success was measured by better real-world decisions, not only offline model metrics.
  • Have strong hands-on experience across the ML lifecycle: dataset design, training or adaptation, evaluation, deployment, monitoring, and iteration.
  • Bring depth in one or more of LLM and retrieval systems, ranking or recommendation, forecasting or anomaly detection, causal ML or experiment analysis, or simulation. You do not need to have done all of them.
  • Have strong software engineering fundamentals and can build high-quality production systems in Python while working comfortably across data, backend, and platform boundaries.
  • Have a strong grounding in machine learning, statistics, computer science, or a related field through formal study or equivalent practical experience.
  • Understand experimentation and statistical reasoning, especially why predictive accuracy is not the same as causal validity.
  • Treat calibration, uncertainty, provenance, privacy, and human review as product requirements, not cleanup work.
  • Can translate ambiguous partner questions into a product and technical roadmap, and work well with product, data science, research, and infrastructure partners.
  • Enjoy building for internal power users and agents, and can make sophisticated ML capabilities feel clear and actionable.
  • Value in-person collaboration and want to help shape a growing Bellevue-based team.

About the company

The Statsig team within OpenAI builds the experimentation, feature rollout, dynamic configuration, and analytics systems that help OpenAI ship products with speed, safety, and evidence. Our work sits on the critical path for how product, engineering, research, and go-to-market teams learn from real-world usage and make high-confidence decisions.

Statsig began as an independent company focused on helping builders move faster through trustworthy experimentation and feature management. After Statsig joined OpenAI, the team began the next chapter: bringing that deep product expertise, customer intuition, and mature platform infrastructure into OpenAI as the experimentation and rollout platform for every product we ship.

Today, we support teams across ChatGPT, Codex, model measurement, consumer experiences, business subscriptions, developer products, and the shared infrastructure that connects them. These teams rely on Statsig to safely introduce new capabilities, compare product and model behavior, measure impact, and roll changes forward or back with confidence.

We are at a defining moment in the platform journey. OpenAI has the data, product surface area, and pace of innovation to learn faster than almost any organization in the world, but that potential only becomes real if teams can experiment responsibly, measure clearly, and roll out changes safely.

We are evolving experimentation systems to help teams learn from product behavior and make better evidence-based decisions.

Based out of OpenAI’s Bellevue office, we are a close-knit team that values in-person collaboration, urgency, craft, and impact. We build for other builders, and the best version of this team is one where every OpenAI product team can move faster because the experimentation and rollout layer is dependable, fast, and easy to use., OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity., At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

7:10 min

Exploring pathways into the machine learning engineering field

Jose Luis Latorre Millas · LIVE

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:37 min

Tracing the evolution from early AI to generative AI

Mike Mike · World Congress 2025

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

Videos

See all

Related articles

See all