About This Session
Internet-scale datasets have successfully driven the evolution of Large Language Models (LLMs) and Vision Language Models (VLMs) across applications like coding assistants and image understanding, Physical AI presents a unique data bottleneck. Physical AI relies heavily on grounded data from sensors, environments, human demonstrations, and real-world simulations. Because this data is often costly, safety-critical, domain-specific, and fragmented, it introduces significant obstacles to model generalization and reliable deployment. This talk addresses the core data challenges in Physical AI, including the scarcity of high-quality embodied datasets, the sim-to-real gap, and the difficulty of capturing rare, long-tail physical scenarios. Finally, we will examine how World Foundation Models can be leveraged to scale data generation and overcome these barriers.
Topics
- Robotics
- Synthetic Data