Researchers introduce Hy-Embodied-RxBrain, a foundation model combining language-visual reasoning and imagination for agents to connect high-level task reasoning with achievable physical states.
The model unifies language and visual imagination in a single planning sequence: language provides abstract plan structure including task decomposition, constraints, and temporal logic, while visual imagination grounds this through world state prediction and joint subgoals.
The model's integration of language planning and visual world modeling addresses embodied cognition challenges relevant to autonomous systems requiring coordinated reasoning and physical control.
Researchers led by Yuan Gao et al. released Chat2Scenic, an iterative RAG-based framework for automatically generating regulation-compliant test scenarios for autonomous driving validation in simulation environments (arXiv 2607.14387, submitted July 2026).
The framework integrates Retrieval-Augmented Generation with a chatbot interface to convert regulatory descriptions into Domain Specific Language (DSL) scenario scripts, addressing the trade-off between compilation success rates and scalability that previous retrieval or retrieval-assemble methods faced.
The authors released an open benchmark comprising 123 scenarios sourced from NHTSA and United Nations Vehicle Regulations, providing the first publicly available dataset for this task and demonstrating significant improvements in automated scenario generation for autonomous driving validation.