A Very Big Video Reasoning Suite

Multimodal Intelligence LVTU VideoQA
视频模型的快速发展长期以来主要聚焦于视觉质量的提升,而对其推理能力的探索却相对不足。视频推理将智能建立在时空一致的视觉环境基础之上,这种环境所蕴含的信息远超文本所能自然表达的范畴,从而支持对连续性、交互性与因果性等时空结构进行直观推理。然而,由于缺乏大规模训练数据,视频推理能力及其随模型规模扩展而表现出的规律性(即“缩放行为”)一直难以开展系统性研究。为填补这一空白,我们推出了“超大规模视频推理数据集”(VBVR),这是迄今规模空前的视频推理资源:它涵盖200项经过精心设计、遵循严谨分类体系的推理任务,并包含逾一百万段视频片段——其体量比现有同类数据集高出约三个数量级。此外,我们还提出了VBVR-Bench评估框架,该框架突破了依赖大模型打分的传统范式,转而采用基于规则、且与人类判断高度一致的评分器,从而实现对视频推理能力可复现、可解释的精准诊断。依托VBVR整套工具,我们开展了迄今首批大规模视频推理缩放研究之一,并首次观察到模型在未见过的新型推理任务上展现出初步的“涌现式泛化”能力。综上,VBVR为构建具备通用性的视频推理能力奠定了坚实基础。全部数据、基准测试工具包及预训练模型均已开源,公众可通过 https://video-reason.com/ 免费获取。
Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causality. However, systematically studying video reasoning and its scaling behavior is hindered by the lack of large-scale training data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks following a principled taxonomy and over one million video clips, approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, benchmark toolkit, and models are publicly available at https://video-reason.com/ .
许愿