Temporal-Distance JEPA: Learning Plan-Aware Latent World Models from Reward-Free Demonstrations时间距离JEPA:从无奖励演示日志学习规划感知的潜在世界模型
S 2.8T11 sources1 个来源R7-research cross-source×2
Temporal-Distance JEPA (TD-JEPA) targets a mismatch in Joint-Embedding Predictive Architectures (JEPAs): JEPA-style training optimizes only short-horizon latent prediction, while planning needs a multi-step ranking of imagined futures by goal progress — a ranking prior JEPA planners typically borrowed from latent Euclidean distance in embedding geometry.
Built on the existing LeWM encoder-predictor and SIGReg backbone, TD-JEPA mines a directed temporal cost directly from reward-free offline demonstration logs, using same-trajectory step order as positive targets, cross-trajectory pairs as heuristic negatives, and a rollout-consistency term to align the learned cost with the planner's horizon.
For autonomous driving, this offers a path to train latent model-predictive controllers directly from recorded driving demonstration logs, supporting trajectory planning and vehicle control without pixel-level reconstruction.