Omniview-Tuning: Boosting Viewpoint Invariance of Vision-Language Pre-training Models

2024年04月18日
Vision-Language Pre-training(VLP)模型,如CLIP,在计算机视觉方面取得了显著成功,特别是在2D图像分布转移方面表现出卓越的鲁棒性。然而,在3D视角变化下,它们的鲁棒性仍然有限,这可能会阻碍实际应用的发展。本文成功地解决了这个问题,同时保持了VLP的原始性能,通过突破两个主要障碍:1)训练数据的稀缺性和2)次优的微调范式。为了解决数据稀缺性,我们构建了Multi-View Caption(MVCap)数据集——一个包含100K多个对象的超过四百万个多视角图像-文本对的全面收集,为VLP模型提供了更多的潜力,以开发可推广的视角不变表示。为了解决现有范式在性能权衡和训练效率方面的限制,我们设计了一种新的微调框架,名为Omniview-Tuning(OVT)。具体而言,OVT通过一种极小极大优化策略引入了交叉视角对齐目标,有效地对齐来自不同视角的相同对象的表示,而不会导致过度拟合。此外,OVT以参数高效的方式微调VLP模型,从而导致最小的计算成本。在各种具有不同架构的VLP模型上进行的广泛实验验证了OVT显着提高了模型对视角转移的鲁棒性,并保持了原始性能,为提高VLP模型的视角不变性建立了先驱性标准。
Vision-Language Pre-training (VLP) models like CLIP have achieved remarkable success in computer vision and particularly demonstrated superior robustness to distribution shifts of 2D images. However, their robustness under 3D viewpoint variations is still limited, which can hinder the development for real-world applications. This paper successfully addresses this concern while keeping VLPs' original performance by breaking through two primary obstacles: 1) the scarcity of training data and 2) the suboptimal fine-tuning paradigms. To combat data scarcity, we build the Multi-View Caption (MVCap) dataset -- a comprehensive collection of over four million multi-view image-text pairs across more than 100K objects, providing more potential for VLP models to develop generalizable viewpoint-invariant representations. To address the limitations of existing paradigms in performance trade-offs and training efficiency, we design a novel fine-tuning framework named Omniview-Tuning (OVT). Specifically, OVT introduces a Cross-Viewpoint Alignment objective through a minimax-like optimization strategy, which effectively aligns representations of identical objects from diverse viewpoints without causing overfitting. Additionally, OVT fine-tunes VLP models in a parameter-efficient manner, leading to minimal computational cost. Extensive experiments on various VLP models with different architectures validate that OVT significantly improves the models' resilience to viewpoint shifts and keeps the original performance, establishing a pioneering standard for boosting the viewpoint invariance of VLP models.
许愿