Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models

Embodied AI and Robotics Task Decomposition RP Multimodal Intelligence Vision-Language Pre-training Visual Grounding
2024年07月09日
近年来,视觉语言导航(VLN)受到越来越多的关注,许多方法已经出现以推进其发展。基础模型的显著成就已经塑造了VLN研究的挑战和提出的方法。在本调查中,我们提供了一个自上而下的回顾,采用了一个基于原则的行动规划和推理框架,并强调了当前方法和未来机会,利用基础模型来解决VLN挑战。我们希望我们的深入讨论可以提供有价值的资源和见解:一方面,里程碑式地记录进展并探索基础模型在该领域中的机会和潜在角色,另一方面,将VLN中的不同挑战和解决方案组织起来,以供基础模型研究人员参考。
Vision-and-Language Navigation (VLN) has gained increasing attention over recent years and many approaches have emerged to advance their development. The remarkable achievements of foundation models have shaped the challenges and proposed methods for VLN research. In this survey, we provide a top-down review that adopts a principled framework for embodied planning and reasoning, and emphasizes the current methods and future opportunities leveraging foundation models to address VLN challenges. We hope our in-depth discussions could provide valuable resources and insights: on one hand, to milestone the progress and explore opportunities and potential roles for foundation models in this field, and on the other, to organize different challenges and solutions in VLN to foundation model researchers.
许愿