FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model

Embodied AI and Robotics Imitation Learning SRT Task Decomposition MLMRD
2026年07月01日
当前关于机器人家具装配的研究大多集中于玩具尺度场景或单臂操作任务。本文提出了“FurnitureVLA”——首个面向真实尺度、双臂协同的家具装配系统性研究,其核心是基于视觉–语言–动作模型(Vision-Language-Action, VLA)构建的完整框架。我们对任务进行了形式化定义,开发了一套可扩展的仿真流水线,用于专家数据的生成与性能评估;同时构建了一套虚拟现实(VR)遥操作系统,支持单操作员以双臂协同方式实施精细控制,从而采集高质量的真实世界装配演示数据。针对极端长时程装配任务(最多包含7个子任务、1550个控制步),我们提出一种“进度增强型VLA”模型:该模型在语义明确划分的子任务上进行微调,可联合预测控制动作与连续的进度信号,从而实现子任务的自动切换,并在推理过程中显著降低误差累积效应。此外,我们深入探究了影响真实尺度装配精度的关键感知与控制设计因素。实验结果表明,在三种不同类型的家具装配任务中,FurnitureVLA相较基线方法将仿真环境下的平均成功率从48%提升至80%;其中,设计因素研究本身进一步带来了21%的成功率增益。我们在真实的Kinova Gen3机器人平台上完成了验证,在最具挑战性的任务中,性能仅下降16%。
Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduce FurnitureVLA, the first systematic study of real-scale bimanual furniture assembly using Vision-Language-Action models (VLAs). We formalize the task, develop a scalable simulation pipeline for expert data generation and evaluation, and build a VR teleoperation system for single-operator bimanual control to collect high-quality real-world demonstrations. To address extreme long-horizon assembly with up to 7 subtasks and 1550 control steps, we propose a progress-enhanced VLA, finetuned on semantically grounded subtasks, that jointly predicts actions and a continuous progress signal, enabling automatic subtask transitions and reducing compounding errors during inference. We further study perception and control design factors that critically affect precision in real-scale assembly. FurnitureVLA improves average simulation success from 48% to 80% compared to baselines across three furniture types, with an additional 21% gain from our design factor study. We validate on a real Kinova Gen3 platform with only 16% drop on the hardest task.
许愿