StreamAV-Bench: A Comprehensive Benchmark for Streaming Audio-Video Generation

Multimodal Intelligence AVEL Other MMI GenAI TDIG Other GenAI
2026年08月26日
生成模型的最新进展正推动视频生成技术迈向面向实时交互式世界的无界流式音视频生成。然而,现有基准测试主要针对已完成序列进行评估,难以准确刻画流式生成所特有的动态特性。为弥合这一差距,我们推出了StreamAV-Bench——首个专为流式音视频生成设计的综合性基准测试。StreamAV-Bench构建了一套统一的评估框架,包含两大评测轨道:其一为“渐进式轨道”,侧重考察模型对指令的逐步遵循能力及长时序下的生成稳定性;其二为“交互式轨道”,重点评估模型在交互过程中的响应能力、状态保持能力以及状态复用能力。该基准覆盖32个细粒度评估维度,并经领域专家严格验证;基于此,我们对13个具有代表性的系统开展了大规模实证评估。分析结果表明,当前模型在渐进式生成中普遍存在时间漂移问题,在交互式控制过程中则面临响应延迟瓶颈。通过系统性的失效归因分析,我们提炼出若干关键洞见,旨在切实推动原生联合音视频流式生成模型的进一步发展。
Recent advancements in generative models are pushing video generation toward unbounded streaming audio-video generation for real-time interactive worlds. However, existing benchmarks primarily evaluate completed sequences and struggle to capture streaming properties. To bridge this gap, we introduce StreamAV-Bench, the first comprehensive benchmark tailored for streaming audio-video generation. StreamAV-Bench establishes a unified evaluation framework, including the progressive track for instruction adherence and long-horizon stability, and the interactive track for interactive response and state retention and reuse. With expert-verified evaluation cases across 32 fine-grained dimensions, we conduct an extensive evaluation of 13 representative systems. Our analysis reveals that current models suffer from temporal drift in progressive generation and responsiveness bottlenecks during interactive control. Based on a comprehensive failure analysis, we share insights to advance the development of native joint audio-video streaming models.
许愿