Phoebe Wang builds machine learning infrastructure at OpenAI. She works on the systems layer behind production AI, focusing on reliability, evaluation, and the daily workflows that keep models running.
Beyond raw infrastructure, her research dives into explainable AI, semantic anomaly detection, and human-centered machine learning. She pushes teams to treat evaluations as production infrastructure rather than mere benchmarks, helping developers expose failure modes and debug complex agent systems. She has also published CHI research on stakeholder fairness in predictive systems.