Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass

AK's Picks CV 3DGS VSER
多视图3D重建仍然是计算机视觉中的一个核心挑战,特别是在需要跨多种视角提供准确且可扩展表示的应用中。目前领先的方法,如DUSt3R,采用的是基本的成对处理方法,即逐对处理图像,并需要昂贵的全局对齐程序以从多个视图中重建。在本文中,我们提出了一种新的多视图泛化方法——Fast3R(快速3D重建),该方法通过并行处理多个视图,实现了高效且可扩展的3D重建。Fast3R基于Transformer架构,在单次前向传递中处理N幅图像,从而避免了迭代对齐的需求。通过广泛的相机姿态估计和3D重建实验,Fast3R展示了最先进的性能,在推理速度上有显著提升,并减少了误差累积。这些结果确立了Fast3R作为多视图应用的稳健替代方案的地位,在不牺牲重建精度的前提下提供了增强的可扩展性。
Multi-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives. Current leading methods such as DUSt3R employ a fundamentally pairwise approach, processing images in pairs and necessitating costly global alignment procedures to reconstruct from multiple views. In this work, we propose Fast 3D Reconstruction (Fast3R), a novel multi-view generalization to DUSt3R that achieves efficient and scalable 3D reconstruction by processing many views in parallel. Fast3R's Transformer-based architecture forwards N images in a single forward pass, bypassing the need for iterative alignment. Through extensive experiments on camera pose estimation and 3D reconstruction, Fast3R demonstrates state-of-the-art performance, with significant improvements in inference speed and reduced error accumulation. These results establish Fast3R as a robust alternative for multi-view applications, offering enhanced scalability without compromising reconstruction accuracy.
许愿