Towards Generalizable Zero-Shot Manipulation via Translating Human Interaction Plans

Embodied AI and Robotics Imitation Learning SRT
我们的目标是开发出能够通过多样化的操作技能与未知物体进行零样本交互的机器人,并展示了被动的人类视频可以作为学习这种通用机器人的丰富数据来源。与典型的机器人学习方法不同,我们采用一种分解方法,可以利用大规模的人类视频学习人类如何完成所需任务(人类计划),然后将其转化为机器人的具体实现。具体而言,我们学习了一个人类计划预测器,它可以根据当前场景图像和目标图像预测未来的手和物体配置。我们将其与一个翻译模块相结合,该模块学习了一个计划条件的机器人操作策略,并允许以零样本方式遵循人类计划进行通用操作任务。重要的是,虽然计划预测器可以利用大规模的人类视频进行学习,但翻译模块仅需要少量领域内数据,可以推广到训练期间未见过的任务。我们展示了我们学习的系统可以执行超过16种操作技能,可以推广到40个物体,涵盖了100个桌面操作和各种野外操作的真实任务。
We pursue the goal of developing robots that can interact zero-shot with generic unseen objects via a diverse repertoire of manipulation skills and show how passive human videos can serve as a rich source of data for learning such generalist robots. Unlike typical robot learning approaches which directly learn how a robot should act from interaction data, we adopt a factorized approach that can leverage large-scale human videos to learn how a human would accomplish a desired task (a human plan), followed by translating this plan to the robots embodiment. Specifically, we learn a human plan predictor that, given a current image of a scene and a goal image, predicts the future hand and object configurations. We combine this with a translation module that learns a plan-conditioned robot manipulation policy, and allows following humans plans for generic manipulation tasks in a zero-shot manner with no deployment-time training. Importantly, while the plan predictor can leverage large-scale human videos for learning, the translation module only requires a small amount of in-domain data, and can generalize to tasks not seen during training. We show that our learned system can perform over 16 manipulation skills that generalize to 40 objects, encompassing 100 real-world tasks for table-top manipulation and diverse in-the-wild manipulation. https://homangab.github.io/hopman/
许愿