New Paper Names 'Thinking in Video' Paradigm, Questions Whether Video Generators Truly Reason新论文提出"Thinking in Video"范式,质疑视频生成模型能否真正推理
S 1.7T11 sources1 个来源R7-research
A 15-author team led by Yongheng Zhang (with Di Yin, Xing Sun, and 12 others) posted arXiv:2607.17523v1 on July 20, 2026, formally naming an emerging paradigm "Thinking in Video," in which video generative models are used to simulate, predict, and reason about real-world dynamics.
The paper redefines video as a medium for constructing, extending, and verifying causal thought rather than merely an output artifact, treating the generation process itself as a form of reasoning.
The authors caution that this reasoning promise remains unverified, since visually convincing video rollouts may reflect memorized appearances rather than genuine causal understanding — a distinction directly relevant to autonomous driving, where world models must reliably predict real-world causal dynamics rather than just produce plausible-looking video for simulation and testing.
由Yongheng Zhang领衔、包含Di Yin、Xing Sun等在内的15位作者团队于2026年7月20日提交论文arXiv:2607.17523v1,正式将利用视频生成模型模拟、预测并推理现实世界动态的新兴范式命名为"Thinking in Video"。