WorldExam Benchmark Tests 20 World Models on Reactivity, Not Just Visual AppearanceWorldExam基准测试20个世界模型的场景反应能力,而非仅视觉表现
S 4.4T11 sources1 个来源R7-research cross-source×2
Yuxue Yang and 15 co-authors (arXiv:2608.02603, submitted Aug 3, 2026) introduce WorldExam, a hierarchical benchmark that evaluates video-generation world models beyond visual appearance, testing whether they can infer from a scene's state how the world should react and generate plausible consequences not explicitly given in the input.
The benchmark comprises 1,474 test cases across 8 tasks and four evaluation levels (Visual Quality, Control Adherence, Spatial Consistency, World Reactivity), spans camera-, action-, and language-driven model paradigms, and was used to assess 20 representative world models.
For autonomous driving, world models with strong reactivity are essential for simulation and motion planning since they must predict how surroundings respond to a vehicle's actions; WorldExam's results show camera-driven models lead on control adherence while action-driven models better capture dynamic interactions — a distinction directly relevant to selecting simulation backbones for AV development.