Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

AK's Picks Embodied AI and Robotics Task Decomposition MLMRD ML RL
我们提出了Embodied-R1.5,这是一种统一的具身基础模型(Embodied Foundation Model, EFM),其单一架构集成了全面的具身推理能力,涵盖具身认知、任务规划、错误修正与指向定位等多个核心维度,旨在迈向通用物理智能。我们构建了三条自动化数据生成流水线,显著拓展了关键能力所需的数据覆盖范围,由此构建了一个规模超150亿词元(tokens)的大型数据系统;同时设计了一种多任务均衡的强化学习训练范式,以有效缓解异构任务之间存在的冲突问题。此外,我们引入了一种“规划器–定位器–修正器”(Planner-Grounder-Corrector, PGC)闭环框架,使单个模型能够自主执行并持续自我修正,从而可靠地完成长时程复杂任务。尽管参数量仅为80亿,Embodied-R1.5在24项具身视觉语言模型(VLM)基准测试中的16项上达到当前最优水平(SOTA),性能超越Gemini-Robotics-ER-1.5和GPT-5.4等前沿模型。得益于模型内部已深度内化的具身能力,Embodied-R1.5仅需少量数据微调,即可高效适配为具身语言动作模型(VLA),并在4个主流操作任务基准套件上全面超越领先VLA模型(如$π_{0.5}$)。我们进一步开展了大量零样本真实机器人实验,全面验证了该模型在指令跟随、功能可供性定位(affordance grounding)、关节式物体操控以及长时程复杂任务等方面的实际表现,充分展现出其面向物理世界的强大泛化能力。为推动具身基础模型领域的后续研究,我们开源了模型权重、全部训练数据集、训练代码,以及专为具身任务设计的评估框架——EmbodiedEvalKit。
We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointing, within a single architecture toward general physical intelligence. Leveraging three automated data construction pipelines to significantly expand the data coverage of critical capabilities, we build a large-scale data system of over 15B tokens, and design a multi-task balanced RL recipe to alleviate heterogeneous task conflicts. We further introduce a Planner-Grounder-Corrector (PGC) closed-loop framework that enables a single model to autonomously execute and self-correct over long-horizon tasks. With only 8B parameters, Embodied-R1.5 achieves SOTA on 16 out of 24 embodied VLM benchmarks, surpassing leading models like Gemini-Robotics-ER-1.5 and GPT-5.4. Benefiting from the internalized embodied capabilities, Embodied-R1.5 can be fine-tuned into a VLA with only a small amount of data, outperforming leading VLA models like $π_{0.5}$ across 4 popular manipulation benchmark suites. We further conduct extensive zero-shot real-robot experiments, validating performance in instruction following, affordance grounding, articulated object manipulation, and long-horizon complex tasks, demonstrating strong generalization to the physical world. We open-source model weights, datasets, training code, and EmbodiedEvalKit, an evaluation framework tailored for embodied tasks, to facilitate future research in EFMs.
许愿