DM2RM: Dual-Mode Multimodal Ranking for Target Objects and Receptacles Based on Open-Vocabulary Instructions

Embodied AI and Robotics RP MLMRD Multimodal Intelligence Vision-Language Pre-training FGCMA
本研究旨在开发一款家庭服务机器人(DSR),该机器人可根据开放式词汇指令将日常物品搬到指定的家具上。目前很少有现有方法能够处理基于图像检索的移动操作任务和开放式词汇指令,并且大多数方法无法识别目标物品和容器。我们提出了双模式多模态排名模型(DM2RM),该模型基于多模态基础模型,能够使用单个模型检索目标物品和容器的图像。我们引入了一个切换机制,利用模式令牌和通过大型语言模型进行短语识别,以根据预测目标切换嵌入空间。为了评估DM2RM,我们构建了一个新的数据集,包括从数百个建筑规模的环境中收集的实际图像和通过众包收集的带有指称表达式的指令。评估结果表明,DM2RM在图像检索设置中的标准度量方面优于先前的方法。此外,我们展示了DM2RM在标准化的真实世界DSR平台上的应用,包括取物和搬运操作,尽管采用了零样本转移设置,但其任务成功率达到了82%。演示视频、代码和更多材料可在https://kkrr10.github.io/dm2rm/上获得。
In this study, we aim to develop a domestic service robot (DSR) that, guided by open-vocabulary instructions, can carry everyday objects to the specified pieces of furniture. Few existing methods handle mobile manipulation tasks with open-vocabulary instructions in the image retrieval setting, and most do not identify both the target objects and the receptacles. We propose the Dual-Mode Multimodal Ranking model (DM2RM), which enables images of both the target objects and receptacles to be retrieved using a single model based on multimodal foundation models. We introduce a switching mechanism that leverages a mode token and phrase identification via a large language model to switch the embedding space based on the prediction target. To evaluate the DM2RM, we construct a novel dataset including real-world images collected from hundreds of building-scale environments and crowd-sourced instructions with referring expressions. The evaluation results show that the proposed DM2RM outperforms previous approaches in terms of standard metrics in image retrieval settings. Furthermore, we demonstrate the application of the DM2RM on a standardized real-world DSR platform including fetch-and-carry actions, where it achieves a task success rate of 82% despite the zero-shot transfer setting. Demonstration videos, code, and more materials are available at https://kkrr10.github.io/dm2rm/.
许愿