# Sci-VBench：科学领域知识密集型视频生成的评测基准

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-08-10 08:00
- AIHOT 分数：45
- AIHOT 链接：https://aihot.virxact.com/items/cmso2m6ot082rrofwnmm5aq7y
- 原文链接：https://arxiv.org/abs/2608.09873

## AI 摘要

Sci-VBench 基准包含 1,253 个专家标注示例，覆盖自然科学、医疗健康、人文社科与工程四大领域的 60 个主题，要求模型生成具备科学推理能力的时序视频。评测显示，16 个前沿模型在感知质量上接近，但在提示词接地性与科学因果正确性上差异显著，专有与开源模型间存在明显差距。

## 正文

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.
