How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu

Machine Learning Street Talk (MLST)
15 September 2026 25 min
0:00 --:--
Episode Description
The car making a left turn at the start of this episode was never filmed. Cosmos 3 generated it. Ming-Yu Liu, who leads the Cosmos research at NVIDIA, explains how one model can describe a video, generate one, and produce robot actions.He walks Tim through the architecture. A vision language model reasons one token at a time; its weights then initialise a bidirectional diffusion generator for video, audio and action, and a shared temporal position scheme lines up signals that run at different ra

Shared via Hopper