诸如 Seedance2.0 和 Veo3.1 等商业视频生成系统已迅速改进,强化了视频生成器可能正演变为“世界模拟器”的观点。然而,社区仍缺乏一个能直接测试模型是否能够推理观察到的世界应如何随时间演变的基准。我们引入了 WorldReasonBench,它将视频生成评估重新定义为世界状态预测:给定初始状态和动作,模型能否生成一个未来视频,使其状态演化在物理、社会、逻辑和信息层面保持一致?WorldReasonBench 包含 436 个精心策划的测试用例,并配有结构化的真实答案问答标注,涵盖四个推理维度和 22 个子类别。我们采用一种与人类对齐的两部分方法来评估生成的视频:过程感知推理验证使用结构化问答和推理阶段诊断来检测时间与因果层面的失败,而多维质量评估则对推理质量、时间一致性和视觉美感进行评分,用于排序和奖励建模。我们进一步引入了 WorldRewardBench,这是一个包含约 6K 个专家标注对(基于 1.4K 个视频)的偏好基准,支持成对和逐点的奖励模型评估。在当前的视频生成器上,我们的结果揭示了视觉合理性与世界推理之间持续存在的差距:视频可能看起来令人信服,但在动力学、因果关系或信息保存方面却存在失败。我们将发布我们的基准和评估工具包,以支持社区对真正具有世界感知能力的视频生成的研究,地址为 https://github.com/UniX-AI-Lab/WorldReasonBench/。
Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests whether a model can reason about how an observed world should evolve over time. We introduce WorldReasonBench, which reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose state evolution remains physically, socially, logically, and informationally consistent? WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions and 22 subcategories. We evaluate generated videos with a human-aligned two-part methodology: Process-aware Reasoning Verification uses structured QA and reasoning-phase diagnostics to detect temporal and causal failures, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. We further introduce WorldRewardBench, a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos, supporting pair-wise and point-wise reward-model evaluation. Across modern video generators, our results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation. We will release our benchmarks and evaluation toolkit to support community research on genuinely world-aware video generation at https://github.com/UniX-AI-Lab/WorldReasonBench/.