> Markdown version of [/jobs/ext/3066457-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/3066457-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** OpenAI Inc. - **Location:** Bellevue, WA, United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Automated Storage and Retrieval Systems, Python (Programming Language), Machine Learning, Large Language Models, Generative AI, Backend, Information Technology, Machine Learning Operations, GPT - **Published:** September 25, 2026 - **Apply:** https://startup.jobs/machine-learning-engineer-core-experimentation-openai-10190867 ## About the Role * Have led ambiguous 0-to-1 production ML products where success was measured by better real-world decisions, not only offline model metrics. * Have strong hands-on experience across the ML lifecycle: dataset design, training or adaptation, evaluation, deployment, monitoring, and iteration. * Bring depth in one or more of LLM and retrieval systems, ranking or recommendation, forecasting or anomaly detection, causal ML or experiment analysis, or simulation. You do not need to have done all of them. * Have strong software engineering fundamentals and can build high-quality production systems in Python while working comfortably across data, backend, and platform boundaries. * Have a strong grounding in machine learning, statistics, computer science, or a related field through formal study or equivalent practical experience. * Understand experimentation and statistical reasoning, especially why predictive accuracy is not the same as causal validity. * Treat calibration, uncertainty, provenance, privacy, and human review as product requirements, not cleanup work. * Can translate ambiguous partner questions into a product and technical roadmap, and work well with product, data science, research, and infrastructure partners. * Enjoy building for internal power users and agents, and can make sophisticated ML capabilities feel clear and actionable. * Value in-person collaboration and want to help shape a growing Bellevue-based team. ## Description We are looking for a Machine Learning Engineer to lead the technical direction for ML-powered experimentation and insights capabilities. You will build production systems that learn from privacy-protected product and experimentation data to generate evidence-backed insights and support decision-making, and help teams decide which ideas are worth testing live. This is an end-to-end, 0-to-1 role. You will work across ML modeling, retrieval and LLM systems, statistical methods, simulation, data and training pipelines, backend services, and user- and agent-facing product experiences. The hard part is not merely producing a plausible answer. It is making each insight and prediction traceable, calibrated, useful, and safe enough to influence real product decisions. Live experiments remain the source of causal validation. You will design systems that make uncertainty explicit, backtest against historical outcomes, compare predictions with online results, learn from misses, and abstain when the evidence is weak. You will preserve clear review, permission, and approval boundaries as automation becomes more powerful. You will collaborate closely with teams building ChatGPT, Codex, model measurement workflows, consumer products, Growth, business subscription experiences, developer products, and shared infrastructure. You will turn their most important learning and decision problems into general platform capabilities that can support the full company., * Set and execute the technical roadmap for Generative Insights and Predictive Experimentation, from early prototypes through production adoption. * Build cross-experiment learning systems that retrieve and synthesize historical experiments, detect recurring effects and segment behavior, reanalyze prior results when data or methods improve, and generate hypotheses with clear evidence and provenance. * Develop predictive models and simulation workflows, including simulation-based evaluation approaches, to estimate likely impact, affected segments, regression risk, and uncertainty before a full live experiment. * Create high-quality datasets and feature or retrieval pipelines from exposures, events, metrics, experiment metadata, and replay data, with strong lineage, freshness, privacy, and data-quality controls. * Establish rigorous evaluation through offline benchmarks, backtests, calibration, drift monitoring, prediction-to-outcome comparisons, and explicit failure or abstention behavior. * Turn models into durable product, API, and agent workflows that move from an insight to experiment design, approval-gated action, and measured learning. * Partner deeply with data science and product teams on experiment design, causal inference, sequential decision-making, variance reduction, and the boundary between prediction and causal evidence. * Build reliable services and intuitive workflows so sophisticated ML capabilities are understandable and useful to teams making high-stakes product decisions. * Provide technical leadership across engineering, product, data science, and research partners, and raise the bar for production ML quality across the platform. ## Related Videos - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Your imaginations is (no longer) the limit: how Generative AI empowers people to be creative](https://www.wearedevelopers.com/videos/741-your-imaginations-is-no-longer-the-limit-how-generative-ai-empowers-people-to-be-creative) - [Navigating the AI Revolution in Software Development](https://www.wearedevelopers.com/videos/1266-navigating-the-ai-revolution-in-software-development) - [Speak, Code, Deploy: Transforming Developer Experience with Voice Commands](https://www.wearedevelopers.com/videos/1159-speak-code-deploy-transforming-developer-experience-with-voice-commands) - [Building Products in the era of GenAI](https://www.wearedevelopers.com/videos/827-building-products-in-the-era-of-genai) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)