Abstract
World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interventions, and adapting when execution departs from expectations. Although progress is often measured by visual fidelity, their value lies in improving behavior. Before reaching for a cup, a person anticipates its weight and resistance to grasping, shaping the hand before contact. Such anticipation is coarse and rarely pictorial, yet it guides action. This raises a central question: which predictive capabilities improve behavior? Existing surveys, organized by architecture, output modality, or application domain, leave this question implicit. We introduce three progressively stronger capability levels: Plausible models preserve task-relevant temporal, geometric, or physical structure; Controllable models additionally predict how interventions alter that structure; and Actionable models translate predictions into measurable gains in planning, action, learning, evaluation, verification, recovery, or data selection. We complement this hierarchy with a 3 x 4 matrix crossing geometry, physics, and action grounding with improvement loops centered on data, rewards, policies, and the model itself. Using this framework, we survey manipulation, navigation, locomotion, autonomous driving, and general embodied learning, tracing technical progressions, clarifying capability requirements, and examining datasets, benchmarks, and evaluation protocols. We identify challenges in long-horizon consistency, uncertainty calibration, causal intervention testing, latency, verification and recovery, and cross-embodiment transfer. This perspective shifts evaluation from visual plausibility toward whether predictions capture task-relevant state, reflect intervention effects, and improve the closed-loop behavior of embodied agents.
Chinese Translation
世界模型通过维护隐状态、预判后果、比较干预效果并在执行偏离预期时进行调整,从而连接具身智能中的感知与决策。尽管其进展常以视觉保真度来衡量,但其真正价值在于改进行为。在伸手拿杯子之前,人会预先估计其重量和抓握阻力,并在接触之前就调整好手的形态。这种预判是粗略的且很少以图像形式呈现,却能有效引导行动。这引出一个核心问题:哪些预测能力真正能改进行为?现有的综述按架构、输出模态或应用领域组织,对这一问题未能给出明确回答。我们提出了三个逐级增强的能力层次:可信模型保留与任务相关的时序、几何或物理结构;可控模型进一步预测干预如何改变该结构;可行动模型将预测转化为在规划、行动、学习、评估、验证、恢复或数据选择方面可度量的收益。我们在这一层级体系之上补充了一个 3×4 矩阵,将几何、物理和动作落地三个维度,与以数据、奖励、策略和模型自身为核心的四类改进环路相交叉。基于这一框架,我们综述了操作、导航、运动控制、自动驾驶和通用具身学习领域,梳理技术演进脉络,阐明能力需求,并考察相关数据集、基准和评估协议。我们识别了长时程一致性、不确定性校准、因果干预测试、延迟、验证与恢复以及跨具身迁移等方面的挑战。本视角将评估重心从视觉可信性转向预测是否能捕捉任务相关状态、反映干预效果,并改善具身智能体的闭环行为。