PAVXploreRL Proposes Action-Exploration RL to Improve PAV World ModelsPAVXploreRL提出动作探索强化学习方法改进PAV世界模型
S 1.7T11 sources1 个来源R7-research
Han Wang, Zijun Wang, Shuoshuo Xue, Rui Cao, Fenjiao Cheng, Xiaodang Liang, and Roy Ka-Wei Lee submitted "PAVXploreRL" (arXiv:2607.16602v1) on 18 Jul 2026, presenting a reinforcement learning framework for action-conditioned world models that act as scalable policy evaluators to reduce reliance on expensive real-world rollouts in embodied AI.
The framework is designed to jointly satisfy three objectives collectively termed PAV — Physical Plausibility (P), Action Adherence (A), and Visual Fidelity (V) — while remaining robust to both in-distribution (ID) expert demonstrations and out-of-distribution (OOD) actions, a combination the authors say existing world models fail to achieve.
Because action-conditioned world models double as policy evaluators, improvements on the PAV criteria are directly relevant to autonomous driving, where they could support more reliable simulation-based validation of vehicle behavior policies under both expert and OOD action scenarios without costly real-world testing.
Han Wang、Zijun Wang、Shuoshuo Xue、Rui Cao、Fenjiao Cheng、Xiaodang Liang和Roy Ka-Wei Lee于2026年7月18日提交论文《PAVXploreRL》(arXiv:2607.16602v1),提出一种面向动作条件世界模型的强化学习框架,该模型作为可扩展的策略评估器,用以减少对昂贵真实世界试验的依赖。