World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy

Embodied AI and Robotics Imitation Learning World Models ML RL
2026年02月06日
近期,机器人世界模型的研究进展借助视频扩散变换器(video diffusion transformers),实现了基于历史状态与动作对未来观测结果的预测。尽管这类模型能够模拟出逼真的视觉效果,但其动作执行精度往往较差,从而制约了其在下游机器人学习任务中的实际应用价值。本文提出一种名为 World-VLA-Loop 的闭环框架,用于联合优化世界模型与视觉–语言–动作(Vision-Language-Action, VLA)策略。我们设计了一种具备状态感知能力的视频世界模型,该模型通过同步预测未来观测结果与奖励信号,充当高保真、可交互的仿真环境。为提升模型可靠性,我们构建了 SANS 数据集——该数据集纳入大量“近成功”轨迹(near-success trajectories),以显著改善世界模型中动作与结果之间的对齐精度。本框架支持在纯虚拟环境中,对 VLA 策略开展强化学习(Reinforcement Learning, RL)的训后优化,全程无需真实物理交互。尤为关键的是,我们的方法构建了一个协同演化的闭环:VLA 策略生成的失败轨迹被持续反馈至世界模型,用以迭代提升其建模精度;而更精准的世界模型又进一步推动后续 RL 优化过程的性能提升。在仿真环境与真实世界任务上的系统性评测表明,该框架仅需极少的真实物理交互,即可显著提升 VLA 策略的整体性能,从而在通用型机器人系统中建立起世界建模与策略学习之间相互促进、共同演进的良性关系。项目主页:https://showlab.github.io/World-VLA-Loop/
Recent progress in robotic world models has leveraged video diffusion transformers to predict future observations conditioned on historical states and actions. While these models can simulate realistic visual outcomes, they often exhibit poor action-following precision, hindering their utility for downstream robotic learning. In this work, we introduce World-VLA-Loop, a closed-loop framework for the joint refinement of world models and Vision-Language-Action (VLA) policies. We propose a state-aware video world model that functions as a high-fidelity interactive simulator by jointly predicting future observations and reward signals. To enhance reliability, we introduce the SANS dataset, which incorporates near-success trajectories to improve action-outcome alignment within the world model. This framework enables a closed-loop for reinforcement learning (RL) post-training of VLA policies entirely within a virtual environment. Crucially, our approach facilitates a co-evolving cycle: failure rollouts generated by the VLA policy are iteratively fed back to refine the world model precision, which in turn enhances subsequent RL optimization. Evaluations across simulation and real-world tasks demonstrate that our framework significantly boosts VLA performance with minimal physical interaction, establishing a mutually beneficial relationship between world modeling and policy learning for general-purpose robotics. Project page: https://showlab.github.io/World-VLA-Loop/.
许愿