LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning

Embodied AI and Robotics Robust MPC Imitation Learning ML RL PGPO
机器人基础模型需要在复杂视觉场景中进行推理,才能在动态环境中执行自适应动作。尽管近期关于潜在推理型视觉-语言-动作(VLA)模型的研究已展现出对细粒度物理动力学的建模能力,但其应用仍主要集中于静态模仿学习范式,严重制约了模型的适应性与泛化能力。本文提出LaST-R1——一种全新的强化学习(RL)后训练框架,旨在高效利用“先潜在推理、再执行动作”的策略。具体而言,我们设计了核心强化学习算法“潜在到动作策略优化”(LAPO),该算法联合优化潜在推理过程与动作生成过程。通过将潜在的思维链(Chain-of-Thought, CoT)推理显式嵌入强化学习优化循环之中,LAPO激发了对物理世界的深层建模能力,从而支撑智能体在交互式环境中实现鲁棒的动作执行。此外,我们还引入了一种自适应潜在思维链机制,使策略能够根据环境状态的多样性,动态调节其推理步长(即推理的深度或广度)。实验表明,在仅需一次示例监督预热(one-shot supervised warm-up)的前提下,LaST-R1在LIBERO基准测试中取得了高达99.9%的平均任务成功率,显著加快了收敛速度,并大幅超越了此前最优(SOTA)方法的性能表现。在真实世界部署中,LaST-R1在涵盖单臂与双臂操作的四项复杂任务上,相较当前最优的监督微调方法,平均任务成功率提升了22.5%。最后,LaST-R1在仿真环境与真实世界环境之间展现出优异的跨域泛化能力。
Robotic foundation models require reasoning over complex visual scenes to execute adaptive actions in dynamic environments. While recent studies on latent-reasoning Vision-Language-Action (VLA) models have demonstrated the capability to capture fine-grained physical dynamics, they remain predominantly confined to static imitation learning, severely limiting their adaptability and generalization. In this paper, we present LaST-R1, a novel reinforcement learning (RL) post-training framework designed to effectively harness "latent reasoning-before-acting" policies. Specifically, we propose Latent-to-Action Policy Optimization (LAPO), a core RL algorithm that jointly optimizes the latent reasoning process and the action generation. By explicitly embedding latent Chain-of-Thought (CoT) reasoning directly within the RL optimization loop, LAPO stimulates profound physical world modeling, which in turn drives robust execution in interactive environments. Furthermore, an adaptive latent CoT mechanism is introduced, allowing the policy to dynamically modulate its reasoning horizon based on diverse environment states. Experiments show that LaST-R1 achieves a near-perfect 99.9% average success rate on the LIBERO benchmark with only one-shot supervised warm-up, significantly improving convergence speed and performance over prior state-of-the-art (SOTA) methods. In real-world deployments, LaST-R1 yields up to a 22.5% average improvement over SOTA supervised fine-tuning approach across four complex tasks, including both single-arm and dual-arm settings. Finally, LaST-R1 demonstrates strong generalization across simulated and real-world environments.
许愿