FactorJEPA: A World Model That Splits Layout, Agents, and Interactions for Chaotic Global South Urban ScenesFactorJEPA:将世界模型分解为布局、主体、交互三通道,聚焦全球南方拥挤城市场景
S 3.3T11 sources1 个来源R7-research
Researchers propose FactorJEPA, a Joint Embedding Predictive Architecture (JEPA) built for DENSEWORLD—populous, crowded, chaotic Global South urban environments—paired with a dataset of 1,000 hours of drive-through, walk-through, and aerial video collected across 22 cities.
Standard JEPA world models degrade under high agent heterogeneity and partial observability; FactorJEPA instead factorizes scene structure into separate layout, agent, and interaction channels via visibility gating and disentangled subspaces to preserve dense interaction dynamics.
World models underpin autonomous driving perception and prediction; this work targets exactly the mixed-traffic, soft-boundary urban conditions that autonomous vehicles must handle to operate in Global South markets.