RxBrain: Foundation Model for Embodied Cognition with Language-Visual Reasoning and ImaginationRxBrain:融合语言-视觉推理与想象的具身认知基础模型
S 2.8T12 sources2 个来源R7-research cross-source×2
Researchers introduce Hy-Embodied-RxBrain, a foundation model combining language-visual reasoning and imagination for agents to connect high-level task reasoning with achievable physical states.
The model unifies language and visual imagination in a single planning sequence: language provides abstract plan structure including task decomposition, constraints, and temporal logic, while visual imagination grounds this through world state prediction and joint subgoals.
The model's integration of language planning and visual world modeling addresses embodied cognition challenges relevant to autonomous systems requiring coordinated reasoning and physical control.