VGGT: Visual Geometry Grounded Transformer

Award AK's Picks CV Monocular or Binocular Depth Estimation and Occlusion Handling VSER
2025年03月14日
我们提出了 VGGT,这是一种前馈神经网络,能够直接从一个、几个或数百个视图中推断出场景的所有关键 3D 属性,包括相机参数、点地图、深度地图和 3D 点轨迹。这种方法在 3D 计算机视觉领域迈出了重要的一步,以往的模型通常受限于单一任务并专门针对这些任务设计。VGGT 方法还具有简单高效的特点,能够在不到一秒的时间内重建图像,并且其效果仍然优于需要通过视觉几何优化技术进行后处理的替代方法。该网络在多个 3D 任务中取得了最先进的结果,包括相机参数估计、多视角深度估计、稠密点云重建和 3D 点跟踪。我们还展示了使用预训练的 VGGT 作为特征提取骨干可以显著提升下游任务的效果,例如非刚性点跟踪和前馈式新视角合成。代码和模型已在 https://github.com/facebookresearch/vggt 公开提供。
We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. This approach is a step forward in 3D computer vision, where models have typically been constrained to and specialized for single tasks. It is also simple and efficient, reconstructing images in under one second, and still outperforming alternatives that require post-processing with visual geometry optimization techniques. The network achieves state-of-the-art results in multiple 3D tasks, including camera parameter estimation, multi-view depth estimation, dense point cloud reconstruction, and 3D point tracking. We also show that using pretrained VGGT as a feature backbone significantly enhances downstream tasks, such as non-rigid point tracking and feed-forward novel view synthesis. Code and models are publicly available at https://github.com/facebookresearch/vggt.
许愿