内容
精选全部 AI 动态AI 日报主题收藏
接入
Agent 接入
更多
关于更新日志反馈
京ICP备2026012723号-2
原文
HuggingFace Daily Papers(社区热门论文)
57

FilmBench:面向电影级视频生成的专业基准测试

2026-07-27 08:00· 1天前
跳到正文
AI 摘要

北京电影学院与华景数字传媒联合推出FilmBench,一个基于电影语言专业标准的文生视频(T2V)与参考视频(R2V)评测基准。其提示词源自20种电影类型的获奖影片片段,评估体系包含3个主轴、12个组件和35项(T2V)+3项(R2V)子指标。在9个T2V和7个R2V模型上的测试显示,模型得分远低于现有网络风格基准,且存在动态美学与多镜头场景下的显著性能差距。

原文 · 未翻译

Shengyi Wang

Niantong Li

Guangzheng Hu

Hong Qi

Fei Ding

Weixu Qiao

Jinlin Wang

Xiaotong Lv

Peng Han

Zimeng Li

Fanshu Ding

Yushu Wang

Han Wu

Jingjing Chen

Chongxiao Wang

Yanhao Wu

Chenglong Huang

Xiaoqian Zhu

Jie Tian

Hua Li

Jingjing Fan

Mingshuang Tang

Zhong Li

Hengxia Qiang

Weibin Chen

Jinyang Zhen

Bing Zhao

Jing Li

Hu Wei

Moku Lab, Hujing Digital Media & Entertainment Group

Beijing Film Academy

qihong@bfa.edu.cn

kongwang@alibaba-inc.com

lj225205@alibaba-inc.com

Abstract

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film-academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman (T2V) and (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.

HuggingFace Daily Papers(社区热门论文)
57导出 Markdown

FilmBench:面向电影级视频生成的专业基准测试

2026-07-27 08:00·1天前
阅读原文· arxiv.org
AI 摘要

北京电影学院与华景数字传媒联合推出FilmBench,一个基于电影语言专业标准的文生视频(T2V)与参考视频(R2V)评测基准。其提示词源自20种电影类型的获奖影片片段,评估体系包含3个主轴、12个组件和35项(T2V)+3项(R2V)子指标。在9个T2V和7个R2V模型上的测试显示,模型得分远低于现有网络风格基准,且存在动态美学与多镜头场景下的显著性能差距。

原文 · 保持原样,未翻译

Shengyi Wang

Niantong Li

Guangzheng Hu

Hong Qi

Fei Ding

Weixu Qiao

Jinlin Wang

Xiaotong Lv

Peng Han

Zimeng Li

Fanshu Ding

Yushu Wang

Han Wu

Jingjing Chen

Chongxiao Wang

Yanhao Wu

Chenglong Huang

Xiaoqian Zhu

Jie Tian

Hua Li

Jingjing Fan

Mingshuang Tang

Zhong Li

1 Introduction

Figure 1: FilmBench evaluation taxonomy: 3 L1 axes, 12 L2 components, 35+3 (R2V-only) L3 sub-metrics, built from clips across 20 film genres curated by expert directors.

Figure 2: The director-driven reverse-engineering pipeline that turns award-winning film clips into film-grade prompts; see §3.1 for details.

Video generation has progressed at an unprecedented pace. Diffusion- and DiT-based models (Veo [4, 24], Kling [12], Seedance [20], Hailuo [17], HappyHorse [6, 7] and others) now produce 15-second photorealistic clips with controllable camera and even synchronized audio, blurring the line between AI-generated footage and professionally produced cinema. A natural next question is whether the capability of these models has reached the pass mark of professional cinematic creation. Modern film production follows a rigorous Cinematic Language (shot scale, camera movement, shot perspective, lighting, scene blocking, composition and film-grade audio design) refined over more than a century of practice. Existing benchmarks were not designed with this professional, academy-taught Cinematic Language in mind, and so far they paint only a coarse, optimistic picture of model capability.

The current benchmarking landscape, and what is missing.

The community has produced a rich body of evaluation suites in the past three years. Foundational multi-dimensional benchmarks such as VBench [9], VBench-2.0 [29] and Video-Bench [5] dissect quality into 9–18 generic dimensions, and learned scorers such as VideoScore [8] push automatic evaluation closer to human preference. Compositional and fine-grained benchmarks (T2V-CompBench [22], FETV [15]) target attribute, motion and interaction binding; temporal, motion and multi-shot benchmarks (ChronoMagic-Bench [27], VMBench [14], SLVMEval [16], MSVBench [21]) probe long-horizon, motion and cross-shot fidelity; reasoning, physical and social benchmarks (TiViBench [2], SVBench [18], WorldJen [10], RBench [3]) stress higher-order or embodied capabilities; the I2V line (ConsistI2V/I2V-Bench [19], UI2V-Bench [28], IP-Bench [13]) focuses on image-conditioned generation; and audio–video and aesthetic benchmarks (AVGen-Bench [30], VGA-Bench [11]) address audio quality and visual aesthetics, respectively.

Figure 3: Fine-grained, per-dimension FilmBench evaluation on a multi-shot reference-scene (R2V) example reverse-engineered from La La Land: Seedance 2.0 (94.97) vs. Vidu Q2 Pro (73.61). Discussed in the text.

Despite this breadth, none of these benchmarks evaluates models against the standards used in real cinematic creation. They share three structural blind spots. First, prompts are sourced from web users, captioners or LLM templates rather than from verified professional shots, so the input distribution drifts away from cinematic content. Second, the implicit content distribution skews toward generic, English, web-style scenes, with no balanced coverage of cinematic genres or culturally diverse film traditions. Third, the evaluation taxonomies leave out cinematic axes such as shot scale, camera movement, lighting, visual style, character performance and audio quality, or collapse them into a single “aesthetic” score. As a consequence, current leaderboards saturate while professional cinematic creators readily reject the same outputs as not film-grade.

FilmBench.

We address this gap with FilmBench, an evaluation benchmark for cinematic text-to-video (T2V) and reference-to-video (R2V) generation, built jointly with directors and faculty from the Beijing Film Academy and the film studio of Hujing Digital Media & Entertainment Group. FilmBench is grounded in three design principles.

Figure 1 gives an overview of the FilmBench evaluation taxonomy. Expert directors from the Beijing Film Academy selected clips spanning 20 film categories—covering the breadth of professional cinematic genres—and used them as the foundation for both the prompt set and the fine-grained evaluation hierarchy. The taxonomy radiates from the three L1 axes through progressively finer L2 components and L3 sub-metrics, while the surrounding film examples illustrate the kinds of cinematic-language challenges each dimension targets. We next describe how these expert-curated clips are turned into structured, film-grade prompts.

(P1) Reverse-engineered from real films (Figure 2): instead of authoring prompts in the abstract, professional directors select clips from award-winning films that stress Cinematic Language and visual expression, reverse-infer draft prompts with a strong multimodal model (Gemini 3.1 Pro), and refine them into structured shot-level prompts that explicitly encode scene, role, prop, shot scale, camera movement, composition, lighting, dialogue and performance, so that every prompt is anchored to a verified professional reference and every fine-grained criterion in the taxonomy is covered. Because these prompts follow real shot lists, most are multi-shot (1,056 of the 1,169 prompts), unlike the single-clip prompts of prior benchmarks. (P2) Academy-aligned cinematic taxonomy: our evaluation dimensions follow the Cinematic Language teaching system of the Beijing Film Academy, organized as three first-level axes (L1): Instruction Following, Temporal Continuity and Aesthetic Quality. These decompose into 12 second-level components (L2) and 35 third-level sub-metrics (L3). T2V and R2V share this main framework; the R2V task additionally adds a Visual Following L2 component of three L3 sub-metrics (scene-space, character-appearance and prop-reference fidelity), each evaluated against the conditioning reference image, giving R2V 13 L2 and 35+3 L3 (per prompt, only the sub-metric matching its reference type is scored). (P3) Expert-grade automatic evaluator: every sub-metric is scored by an in-house, professional film-grade evaluation agent whose core Cinematic Language operator suite (FilmOps) we open-source; its model-level ranking reproduces expert film-industry rankings at Spearman (T2V) / (R2V) (§3, §4).

A worked example.

Figure 3 shows FilmBench scoring a multi-shot reference-scene (R2V) example, a four-shot clip reverse-engineered from the musical La La Land, on Seedance 2.0 (94.97) and Vidu Q2 Pro (73.61). Because every L3 sub-metric is scored independently, the evaluation exposes behavior that an aggregate score hides. Although Seedance 2.0 leads overall by a wide margin, its scene-space fidelity is slightly lower (87.5 vs. 100): Vidu Q2 Pro over-adheres to the reference background image and scores lower on nearly every other dimension. It drops markedly on Cinematic Language following (shot scale 75 vs. 25, viewing angle 100 vs. 50, camera movement 100 vs. 50, composition 75 vs. 43.75) and, having sacrificed the shot/reverse-shot staging to lock onto the scene, exhibits character-positioning drift that drives its character blocking to 0. FilmBench can thus objectively credit a single dimension (reference fidelity) while still penalizing the broader loss of professional Cinematic Language; Appendix A discusses further per-dimension case studies.

Findings.

We benchmark leading video-generation models (9 for T2V, 7 for R2V), including Seedance 2.0, HappyHorse 1.1/1.0, Kling 3.0, Kling 3.0 Omni, Veo 3.1, Grok Imagine Video [26], Vidu [1], and Hailuo 2.3, and distill five field-wide findings. (i) No saturation: no model approaches the ceiling on either task (top scores of 88.93 for T2V and 86.66 for R2V), in sharp contrast with the near-ceiling numbers on prior web-style benchmarks; and crucially, what separates the leaders is not generic quality but the professional Cinematic Language sub-metrics of Instruction Following (camera movement, shot scale, viewing angle, focus and composition) together with the Aesthetic Quality axis, which is exactly where FilmBench resolves differences that web-style benchmarks miss. (ii) A field-wide dynamic-aesthetics bottleneck: across all models the lowest-scoring sub-metrics are action performance, camera-work appeal, character-motion realism and emotional performance, whereas static image-quality sub-metrics (sharpness, lighting) score markedly higher: today’s models render clean frames but still struggle with believable motion and performance. (iii) Multi-shot is the hardest regime: on T2V, moving from single- to multi-shot prompts costs points on average and up to for the weakest model, concentrated in the Cinematic Language and editing demands of Instruction Following (framing, camera work and cut fluency); it is multi-shot staging, rather than single-take quality, where current models fall short. (iv) Reference conditioning stresses rather than reorders the field: across scene, prop and character references the R2V ranking stays stable, yet conditioning lowers scores to varying degrees and penalizes the weakest model most, and the visual-following leader is not necessarily the overall leader, so reference behavior is best read through ranking shifts rather than absolute per-type scores. (v) No model wins on all L3 sub-metrics: Seedance 2.0, the overall leader, claims only 18/35 (T2V) or 20/38 (R2V) L3 championships, concentrated on Cinematic Language sub-metrics (Cinematic Language, editing appeal and performance), while HappyHorse 1.1 dominates 11/35 (T2V) or 10/38 (R2V) sub-metrics on character, audio and scene dimensions; notably the three R2V-specific visual-following championships all go to the HappyHorse family; 4 (T2V) and 2 (R2V) models do not claim any championship. This “championship mismatch” reveals structural complementarity concealed by a single leaderboard number.

Contributions.

We make three contributions. (i) We release FilmBench, a cinematic-grade T2V/R2V benchmark whose prompts are reverse-engineered from award-winning real films and curated jointly with directors and faculty from the Beijing Film Academy and a professional film studio, ensuring coverage of every fine-grained cinematic criterion. (ii) We release FilmOps, a modular Cinematic Language operator suite with trained weights and inference scripts, enabling the community to build customizable, expert-knowledge-driven judge agents for video evaluation. (iii) We propose an academy-aligned evaluation taxonomy that translates professional Cinematic Language into a three-level hierarchy of 3 L1 axes, 12 L2 components and 35 L3 sub-metrics (35+3 with the R2V-specific Visual-Following extension), scored by an expert-grade automatic evaluator with an open-source Cinematic Language operator suite. (iv) We conduct a comprehensive cinematic evaluation over leading T2V/R2V models, show the automatic evaluator reproduces expert rankings at (T2V) / (R2V), and expose dynamic aesthetics as the shared bottleneck, discussing directions for future research.

2 Related Work

We trace how video-generation evaluation has evolved through three stages (from generic quality scoring, to capability-specific stress tests, to increasingly film-like assessment) and show that each stage, in closing one gap, exposes the next, until the accumulated gaps motivate FilmBench precisely.

From distribution metrics to multi-dimensional, human-aligned evaluation.

Early evaluation repurposed action-recognition datasets (UCF-101, MSR-VTT, Kinetics) and reported single scalars such as FVD/IS, which neither localize which aspect of a video fails nor align well with human perception. As diffusion- and DiT-based generators—including the systems we evaluate (Veo-3.1, Kling-v3/Omni, Seedance-2.0, Hailuo-2.3, Vidu-Q3-Pro, Grok-Imagine-Video and HappyHorse-1.x)—raised the ceiling to controllable, audio-equipped clips, the community moved to disentangled, human-aligned protocols. VBench [9] and VBench-2.0 [29] decompose quality into 16 then 18 dimensions and push from superficial to intrinsic faithfulness; Video-Bench [5] enlists MLLMs as scalable judges; VideoScore [8] learns a human-aligned regressor. This stage established the grammar of modern evaluation—quality as a structured set of interpretable dimensions—but its axes are deliberately generic and its prompts are drawn from web users, reflecting what is easy to source rather than what a film production demands. The first gap is thus one of content and taxonomy: neither the prompts nor the dimensions are anchored in a professional production grammar.

Specializing to emerging capabilities.

With this foundation in place and baseline visual quality saturating, benchmarks specialized to isolate the hard capabilities that generic scores had masked. Compositional suites T2V-CompBench [22] and FETV [15] expose brittle attribute, motion and spatial-relation binding; the temporal–motion line ChronoMagic-Bench [27] and VMBench [14] target metamorphic amplitude and perception-aligned motion, while SLVMEval [16] extends reliability testing to hour-scale footage. As generators began stitching shots into a story, MSVBench [21] introduced multi-shot evaluation with a hybrid LMM-plus-expert scorer, distilling this stage’s lesson: current systems behave as visual interpolators rather than world models. A parallel reasoning wave then raised the bar from rendering to understanding, probing structural (TiViBench [2]), social (SVBench [18]), broad multi-axis (WorldJen [10]) and embodied-robotic (RBench [3]) reasoning. Yet even MSVBench scores its shots against generic quality criteria rather than a director’s intended shot list, so evaluation still stops short of the film: cross-shot narrative structure and director intent are not judged end-to-end against a verified cinematic reference. This is the second gap.

Toward film: reference, audio, aesthetics and Cinematic Language.

The strand closest to our goal relaxes that assumption toward controllable, film-like generation, but does so one aspect at a time. Reference/image-conditioned suites (ConsistI2V/I2V-Bench [19], UI2V-Bench [28] and IP-Bench [13]) evaluate consistency, semantic understanding and protection from a conditioning image, raising reference fidelity as a concern. As generators acquired joint audio, AVGen-Bench [30] evaluates text-to-audio-video generation, adding a cross-modal audio axis, while VGA-Bench [11], co-developed with the Beijing Film Academy, brings aesthetic tagging into view. In parallel, a cinematic-understanding literature (MovieNet, shot-type taxonomies, CineScale/ShotBench) and storyboard-to-video pipelines treat film as structured language, and MovieBench [25] pushes to feature-length identity consistency—sustained in at most of cases at multi-scene scope. Collectively these efforts touch nearly every ingredient of film, yet no single benchmark unifies them: aesthetics and audio are not anchored in a production grammar, reference fidelity and cross-shot audio–dialogue continuity are evaluated in isolation if at all, the target is typically real or arbitrary footage rather than generated films, and the cinematic-understanding line asks only whether models can read Cinematic Language, never whether a generator can speak it at a director’s granularity. Closing this third gap—a unified, production-grounded evaluation of generated film—is exactly the remit of FilmBench.

Positioning of FilmBench.

These three accumulated gaps (a generic taxonomy, a clip-level scope, and fragmented coverage of individual film aspects) jointly define what a film-grade benchmark must satisfy at once. To the best of our knowledge, no existing benchmark simultaneously (i) reverse-engineers prompts from genuine, award-winning films selected by directors; (ii) spans a broad set of cinematic genres and both text-to-video and reference-to-video tasks, with prompts organized as real, mostly multi-shot shot lists (1,056 of 1,169) that evaluate the film rather than the isolated clip; (iii) grounds its taxonomy in the Cinematic Language system taught in professional film schools: 3 first-level axes, 12 second-level components and 35 third-level sub-metrics with an R2V-specific Visual-Following extension, co-designed with directors, Beijing Film Academy faculty and a professional film studio; and (iv) treats audio–dialogue continuity (including dialogue intelligibility and per-character timbre stability across shots) as a first-class axis. FilmBench thus asks not whether a model can read Cinematic Language but whether it can speak it, and because every prompt is reverse-engineered from a verified shot, we always retain both a structured grammar and a ground-truth reference. §3 and §4 detail this taxonomy and the automatic-plus-human evaluator stack, with a head-to-head comparison table in §3 and the appendix.

3 FilmBench: Construction and Dimensions

This section describes how FilmBench is built. We first present the reverse-engineering pipeline (Figure 2) that turns genuine cinematic clips into structured prompts (§3.1), then introduce our three-level Cinematic Language taxonomy of 3 first-level axes (L1), 12 second-level components (L2) and 35 third-level sub-metrics (L3), co-designed with faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio (§3.2).

3.1 Reverse-Engineered Prompt Construction

FilmBench prompts are produced by a director-driven reverse-engineering pipeline with three stages (Figure 2). Clip selection: professional directors from the film studio of Hujing and the Beijing Film Academy prioritize award-winning films (e.g., winners of the Academy Awards, the Golden Horse Awards and the Hundred Flowers Awards) across our 20 cinematic genres, and select multi-shot clips (mostly shots) that specifically stress Cinematic Language and visual-expression skill. Draft prompt inference: a FilmOps operator suite together with a strong multimodal model (Gemini 3.1 Pro) reverse-extracts the original narrative script, Cinematic Language elements and on-screen visual elements from each clip, which Gemini 3.1 Pro then assembles into an initial structured prompt. Expert refinement: the Beijing Film Academy directing team screens, quality-checks, rewrites and standardizes every draft into professional film-grade Cinematic Language. Throughout clip selection and final curation, the team ensures that every L1 axis and L2 component in the taxonomy is covered, so that the benchmark faithfully probes a model’s film-grade generation ability rather than generic visual quality. Each prompt uses slotted cinematic tags (@scene, @role, @prop), e.g., “Shot 1 (medium-shot, eye-level fixed, rule-of-thirds): in a brightly lit ICU @scene1, a young female doctor @role1 ”.

Tasks and scale.

FilmBench covers two tasks: T2V (prompt video) and R2V (reference prompt video). The T2V suite contains 515 prompts spanning 20 base cinematic genres, balanced across Chinese Films and International Films markets (243 Chinese Films / 272 International Films). The R2V suite contains 654 prompts split across three reference types (232 scenes / 209 props / 213 characters) and drawn from the same 20-genre pool across both markets (239 Chinese Films / 415 International Films). Because prompts follow real shot lists, most are multi-shot: 402 of the 515 T2V prompts and all 654 R2V prompts script multiple shots (1,056 of 1,169 overall), unlike the single-clip prompts of prior benchmarks. Each prompt is reverse-engineered from a distinct award-winning film clip, so the prompt count also reflects the number of source clips.

3.2 Academy-Aligned Evaluation Taxonomy

Our evaluation dimensions follow the academic Cinematic Language system taught in professional film schools. The taxonomy is co-designed by faculty from the Beijing Film Academy and the film studio of Hujing together with an analysis of current video-generation models’ capabilities, and is deliberately organized as a multi-level hierarchy so that conclusions and insights can be drawn at coarse (L1), medium (L2) and fine (L3) granularity.

媒体内容 · 前往原文查看

Table 1: FilmBench three-level taxonomy (3 L1 axes, 12 L2 components, 35 L3 sub-metrics). R2V adds a Visual Following component (+3 L3). Full definitions in Appendix B.

L1 Axis L2 Component

Instruction Following (16+3) Cinematic Language

Character & performance

Scene

Visual Following (R2V only)

Temporal Continuity (7) Spatial coherence

Temporal coherence

Character coherence

Audio coherence

Aesthetic Quality (12) Base quality

Editing

Performance

Audio quality

Three L1 axes.

FilmBench evaluates cinematic competence along three first-level axes, decomposed into 12 L2 components and 35+3 L3 sub-metrics (Table 1); below we enumerate the Cinematic Language sub-metrics and leave the remaining, more generic components to Table 1 and Appendix B. Instruction Following captures whether the video faithfully realizes the prompt’s cinematic intent, over four components: Cinematic Language (shot scale, camera movement, viewing angle, composition, focus, tone & color, thematic style), Character & performance, Scene, and Audio. Temporal Continuity captures whether the output behaves as a film rather than a single clip, over four components: spatial coherence, temporal coherence, character coherence, and audio coherence. Aesthetic Quality captures production-grade quality, over four components: base quality, editing (editing fluency, camera-work appeal), performance (emotional performance, action performance), and audio quality.

Task-specific extension.

T2V and R2V share this main framework. R2V additionally introduces one L2 component, Visual Following, with three L3 sub-metrics (scene-space, character-appearance and prop-reference fidelity), each measuring adherence to the conditioning reference image. Every R2V prompt conditions on one of three reference types (scene, prop or character), and only the matching fidelity sub-metric is scored for that prompt. Visual Following is scored only for R2V, so R2V has 13 L2 and 35+3 L3 dimensions, while all other dimensions are identical across tasks. For every L3 sub-metric we provide a 5-point anchor rubric; the complete definition table is in Appendix B.

4 Evaluation Method

FilmBench is scored by an expert-grade automatic evaluation agent which integrates a suite of Cinematic Language operators, expert annotation models, and a judge model. For every dimension in the taxonomy the agent produces a 1–5 score, achieving high agreement with human experts (§4.2). Below we describe the open-sourced operator suite (§4.1) that grounds the Cinematic Language dimensions and the score-aggregation rule (§4.2).

4.1 FilmOps: A Cinematic Language Operator Suite

A key obstacle to professional cinematic evaluation is that neither generic MLLM-as-judge approaches nor existing domain expert models reliably encode film-industry visual rules: general multimodal models are trained on open-domain understanding and misjudge professional categories such as shot scale, composition and camera movement [23], while existing aesthetic/shot models are trained mostly on everyday photography or live-action footage and do not transfer across genres (3D/2D animation, stylized content). We therefore build FilmOps, an open-source operator suite that maps a generated video into structured cinematic labels for the Cinematic Language dimensions.

Industry-aligned taxonomy.

FilmOps defines a professional classification system, aligned with classical production references and vetted by front-line practitioners, over six core dimensions: shot scale, composition, viewing angle, tone & color, character layout, and camera movement. Five are frame-level and one (camera movement) is shot-level; together they span 55 sub-categories (character layout is an open-ended natural-language field). Categories are defined by narrative/visual function rather than imaging mechanism, so the same standard applies across live-action, 3D and 2D genres.

Multi-genre training data.

To match the cross-genre, multi-style distribution of generated video, FilmOps is trained on a pool of 5,000+ real film/TV works spanning live-action, 3D animation, 2D animation and stylized/VFX content, and across narrative types (dialogue-driven, action, group scenes, and long-tail shot forms). Each operator is trained on 40K–60K annotated samples with a shot-based test set, under a strict annotator-training and quality-control protocol led by professional practitioners.

Task-matched operator design.

Operators are built to match each dimension’s characteristics: frame-level visual dimensions (shot scale, composition, angle, color) use vision backbones (DINO, BEiT and ResNet-18), while dimensions requiring relational reasoning or temporal modeling (character layout, camera movement) use a multimodal model (InternVL3-14B, LoRA-SFT). Classification operators are evaluated by macro-F1 and the natural-language layout operator by precision. Against four strong zero-shot general-MLLM baselines, the trained operators improve markedly on every dimension, confirming the value of domain-specialized operators; the full per-operator backbones and macro-F1 figures tested over 400 images/shots against these baselines are reported in the appendix (Table 12).

To support customizable, personalized self-evaluation, we open-source the classification standard, model weights and inference scripts of FilmOps. The modular design is not tied to a single model, so users can reuse individual operators or embed specific dimensions into their own evaluation pipelines.

4.2 Score Aggregation

Each L3 sub-metric receives a 1–5 score, which is linearly mapped to a 0–100 scale before aggregation (). Aggregation is sample-level: each video is first aggregated across its dimensions, and per-model scores are the mean over that model’s videos, preserving sample variance for significance analysis. We aggregate L3 L1 directly (each L1 axis is the mean of its constituent L3 sub-metrics), treating every L3 sub-metric as equally important, and the Overall score is the equal-weighted mean of the three L1 axes; L2 components are computed as a side branch for fine-grained reporting only. Not-applicable scores (N/A or null) are excluded from aggregation; for models without audio capability, audio dimensions are handled symmetrically as not-applicable.

Qualitative case studies.

Fine-grained per-dimension scoring exposes trade-offs that aggregate scores hide. As a worked example, Figure 4 scores a multi-shot T2V sci-fi mech battle, a case that stresses the dual challenge of intense action combined with multi-shot cutting. Seedance 2.0 (86.11) clearly outperforms Grok Imagine Video (55.97), and the gap is diagnostic rather than uniform: Seedance is markedly stronger on tone & color, character & performance and scene, and on camera-work appeal; in the action passages its character positioning stays consistent across shots and its temporal logic shows no visible errors. Grok, by contrast, exhibits scene drift and drifting cross-shot appearance and positioning, which surfaces directly in the per-dimension gaps on character-motion realism and action performance. The introduction (Figure 3) shows a complementary multi-shot R2V reference-scene case, and Appendix A provides two further worked examples (reference-character and reference-prop) that together cover the major scoring scenarios in FilmBench.

Figure 4: Multi-shot T2V example (a sci-fi mech battle): Seedance 2.0 (86.11) vs. Grok Imagine Video (55.97). Discussed in the text.

5 Experiments

5.1 Models and Setup

We benchmark leading video-generation models under each task. For T2V we evaluate 9 models: Seedance 2.0 [20], HappyHorse 1.1 [7], HappyHorse 1.0 [6], Kling 3.0 [12], Kling 3.0 Omni [12], Veo 3.1 [4, 24], Grok Imagine Video [26], Vidu Q3 Pro [1] and Hailuo 2.3 [17]. For R2V we evaluate 7 models: Seedance 2.0, HappyHorse 1.1, HappyHorse 1.0, Kling 3.0 Omni, Veo 3.1, Grok Imagine Video, and Vidu Q2 Pro [1]. All models are run under their official APIs / public checkpoints with default settings, generating one video per prompt. This yields machine-scored T2V records and R2V records. Throughout, T2V and R2V are reported with the same analysis format: overall, per-axis (L1), per-component (L2), per-sub-metric (L3), discriminability, and capability profiles, so that the two tasks are directly comparable; R2V additionally carries a reference-type analysis (scene / character / prop).

5.2 Human Agreement

We validate the automatic evaluator against 90 professional human raters majoring in film, television, directing, or digital media from the Beijing Film Academy at the model level: over 300 prompts (30% of the benchmark) carry both machine and human judgments, yielding model–prompt pairs. We compute per-model machine and human means at each L2 component and correlate the two rankings (Table 2). The overall model-level Spearman reaches (T2V) and (R2V), confirming that the automatic leaderboard closely reproduces expert model rankings.

媒体内容 · 前往原文查看

Table 2: L2 component-level human agreement (Spearman ), averaged across T2V and R2V. 10 of 13 components reach ; the three lower-agreement components (audio quality, audio coherence, editing) are separated by a mid-rule.

Perform.

Vis. Follow.

Cinem. Lang.

Base qual.

Char. & perf.

Temp. coh.

Spat. coh.

Scene

Char. coh.

Audio q.

Audio coh.

Editing

0.99 0.96 0.95 0.95 0.93 0.92 0.85 0.83 0.82 0.78 0.68 0.68 0.60

Drilling into the component level, 10 of the 13 L2 dimensions achieve ; the three lower-agreement components—audio quality (), audio coherence () and editing ()—are those where subjective temporal behaviour (sound fidelity, cut rhythm) is hardest to pin down, and we acknowledge that human and automatic judgments diverge more on these axes. The overall-level agreement ( on both tasks) nonetheless indicates that the aggregate ranking is a reliable proxy for expert judgment, and we flag the three lower-agreement dimensions when interpreting fine-grained results.

5.3 Overall Leaderboard

Tables 7–8 and Figure 5 report the overall FilmBench score (0–100, higher is better). On T2V, Seedance 2.0 leads (88.93), closely followed by HappyHorse 1.1 (87.42) and HappyHorse 1.0 (87.02); the two Kling 3.0 variants form a second tier (86), Vidu Q3 Pro, Grok and Veo 3.1 a third tier (81), with Hailuo 2.3 at 68.94. On R2V, the top group is Seedance 2.0 (86.66), HappyHorse 1.1 (85.51) and HappyHorse 1.0 (84.87). Crucially, no model approaches saturation, in sharp contrast with the near-ceiling scores reported on prior web-style benchmarks, and the top group is consistent across both tasks, indicating that reference conditioning does not reshuffle the leaders.

Figure 5: Overall FilmBench scores (0–100) per model; model names are printed inside each bar and the dashed line is the cross-model average. Left: T2V (); right: R2V ().

5.4 Dimension Variance

To quantify where models differ most at fine-grained levels, we compute the cross-model variance of each dimension across the hierarchy (Figure 6; Hailuo is excluded from the variance computation to avoid its audio outliers dominating). At the axis level, Instruction Following shows by far the largest variance (), nearly eight times that of Aesthetic Quality () and seventeen times that of Temporal Continuity (), indicating that models differ primarily in their ability to execute professional cinematic language instructions.

Drilling into the component level, Cinematic Language dominates with a variance of , far exceeding Audio (), Scene () and all other components. At the sub-metric level, the top five highest-variance dimensions are all Cinematic Language or scene-related: camera movement (), focus (), shot scale (), viewing angle () and fore/mid/background (). This concentration confirms that the primary differentiator among current generative video models lies in Cinematic Language instruction following—whether a model understands and executes professional camera-language conventions. Per-task breakdowns are provided in Appendix C.

Figure 6: Cross-model variance per L3 sub-metric (main bars, colored by axis), with L2 inset (upper right), computed over the merged T2V+R2V model means (Hailuo excluded).

5.5 Per-Axis Landscape (L1)

Across the three axes, models are most differentiated on Instruction Following, which encompasses the Cinematic Language sub-metrics that drive the high variance reported above. On T2V (Figure 7, left), Seedance 2.0 leads all three axes: Instruction Following (91.69, ahead of HappyHorse 1.1 at 90.70), Temporal Continuity (96.06) and Aesthetic Quality (79.03). On R2V (Figure 7, right), Seedance 2.0 likewise leads every axis (Instruction Following 87.45, Temporal Continuity 94.09, Aesthetic Quality 78.44). The Instruction Following axis shows the widest inter-model gap under both tasks, as the Cinematic Language and scene-related sub-metrics amplify the spread in professional-language execution.

Figure 7: Per-axis rankings for T2V (left, 9 models) and R2V (right, 7 models). Each model is scored on three L1 axes: Instruction Following (solid), Temporal Continuity (hatched), and Aesthetic Quality (dotted). Models are ranked by overall score; R2V Instruction Following includes the visual-following component.

5.6 Per-Component Rankings (L2)

Descending one level, we rank models within the L2 components most closely tied to Cinematic Language together with the adjacent Scene and Character & performance components (Figure 8); the full 12/13-panel rankings are deferred to the appendix. Cinematic Language is placed at the centre and used as the sort key, and each component isolates a different craft skill, revealing where the head models specialize:

Cinematic Language (camera work, framing, transitions) is the sharpest discriminator of the three, with the field fanning out over a 17–42-point range and even the strongest model reaching only 84.6. This is the one component that hinges on temporal decision-making—when to move the camera, how to reframe, when to cut—rather than per-frame rendering, which is why models diverge so widely here. Seedance 2.0 leads decisively and is the only model whose camera competence holds up when a reference image is imposed, marking dynamic camera control as its signature strength.

Figure 8: Per-model rankings within the L2 components most related to audiovisual language: Cinematic Language (solid, centre, model names inside; also the sort key), flanked by Scene (hatched) and Character & performance (dotted). T2V (left, 9 models) and R2V (right, 7 models). The full L2 rankings are in the appendix.

Scene (environment consistency, lighting, background layout) also spreads the field, though the leading cluster reaches the mid-90s (HappyHorse 1.0 at 96.0, HappyHorse 1.1 at 96.7). The gap opens further down the ranking, where weaker models struggle to keep a coherent, consistently lit environment across shots. The separation here is owned by the HappyHorse family, whose strength lies specifically in preserving spatial layout—a capability distinct from camera craft, since the models that lead scene rendering are not the ones that lead Cinematic Language.

Character & performance spreads the field the least of the three but still meaningfully, with the leaders reaching the high-90s (HappyHorse 1.1 at 97.6, Seedance 2.0 at 96.3). Here differentiation comes from sustaining believable characters and acting across a full scene, and HappyHorse 1.1 holds a slim but consistent edge, making it the character specialist just as Seedance 2.0 is the camera specialist. Read together, the three components show that the overall leaderboard masks a division of labor—camera craft, scene rendering and character performance are led by different models—so no single system dominates every axis of film craft.

5.7 Fine-Grained Rankings (L3)

At the finest level, FilmBench resolves 35 (T2V) / 38 (R2V) L3 sub-metrics, each with its own ranking; Figure 9 shows the three sub-metrics selected by highest merged T2V+R2V variance. All three are Cinematic Language sub-metrics, indicating that camera craft is where the leading models pull apart most:

Camera movement shows the largest spread among the head models, with scores ranging from Seedance 2.0’s 86.5 down toward the low 50s. Executing a motivated camera move—a push-in, a pan that tracks action, a crane reveal—requires the model to sustain a coherent scene through time, and models diverge considerably in how convincingly they do so. Seedance 2.0 leads by a clear margin, suggesting a training emphasis on dynamic camera choreography that sets it apart from the rest of the field.

Figure 9: Per-model rankings for the three highest-variance L3 sub-metrics (computed over merged T2V+R2V data): Camera movement (hatched), Focus (solid, model names inside), and Shot scale (dotted). T2V (left, 9 models) and R2V (right, 7 models). The full per-metric panels are in the appendix.

Focus (depth-of-field control, subject isolation) sees the leaders cluster in the low-to-mid 80s (Seedance 2.0 at 85.6, HappyHorse 1.1 at 82.7), with a wider tail below. The metric separates models because it rewards intentional focus that directs attention to the dramatically relevant subject—a semantic judgement, beyond merely producing a shallow-depth look, that the stronger models handle more reliably.

Shot scale (close-up, medium, wide framing decisions) tests whether a model chooses the framing a scene calls for—an intimate close-up for a line of dialogue, a wide for an establishing beat. Seedance 2.0 again leads (79.8) with HappyHorse 1.1 the runner-up, and the head ordering across all three sub-metrics is consistent—Seedance 2.0 first, the HappyHorse family close behind. Because these are the dimensions on which models differ most, holding a high standard across them is what distinguishes a film-grade leader: Seedance 2.0’s overall lead comes not from a single capability but from staying near the top on the most challenging Cinematic Language dimensions at once, while the HappyHorse family remains competitive across them.

Crucially, no model wins on all L3 sub-metrics: Seedance 2.0, the overall leader, claims only 18/35 (T2V) or 20/38 (R2V) championships, concentrated on Cinematic Language sub-metrics—Cinematic Language (camera movement, shot scale, focus, viewing angle), camera-work appeal, and performance (action and emotional performance)—in both tasks. HappyHorse 1.1, the runner-up, dominates 11/35 (T2V) or 10/38 (R2V) sub-metrics on character instruction-following, audio, and scene dimensions. Notably, on the three R2V-specific visual-following sub-metrics (scene-space fidelity, character-appearance fidelity and prop-reference fidelity), the HappyHorse family takes all three championships—HappyHorse 1.0 leads scene-space and character-appearance fidelity, HappyHorse 1.1 leads prop-reference fidelity. This “championship mismatch” is the key value of fine-grained evaluation: a single leaderboard number conceals the structural complementarity between models. Beyond the top two, the remaining championships are scattered across several models, with 4 (T2V) or 2 (R2V) models winning no championships at all. This granularity turns a single leaderboard into an actionable diagnostic: a mid-tier model can still lead on individual craft dimensions. The complete per-model score matrices (L3 and L2 heatmaps) are in the appendix (Figures 20–23).

5.8 Market and Genre Robustness

FilmBench spans 20 cinematic genres across Chinese Films and International Films markets (T2V /; R2V /) to ensure broad diversity and coverage. Figure 10 splits each task by market: the top group (Seedance 2.0, HappyHorse 1.1/1.0) is identical for Chinese Films and International Films prompts on both tasks, with only minor mid-pack reordering. Figure 11 shows the score heatmap across the top 4 genres by sample size: rankings remain stable across genres, and each model’s per-genre spread is small (2–3 points), confirming that both leaderboards reflect model capability rather than genre mix. The action subset, often cited as the hardest, shows the same ranking with only a slightly lower ceiling.

Figure 10: Overall ranking split by market, Chinese Films vs. International Films; left: T2V (/), right: R2V (/).

Figure 11: Overall score heatmap across the top 4 genres by sample size. Y-axis: genres (with sample sizes T2V/R2V); X-axis: models split into T2V (left) and R2V (right) sections. Each cell shows the genre-specific mean score.

5.9 Prompt Complexity and Content Factors

Beyond market and genre, we probe three prompt-level factors that stress generation differently: narrative structure (single- vs multi-shot prompts), content type (action vs dialogue scenes), and rendering style (live-action vs animation). In T2V, 113 of 515 prompts are single-shot and 402 are multi-shot; every R2V prompt is multi-shot, so the shot contrast is T2V only. We tag 117 of 515 T2V prompts and 190 of 654 R2V prompts as action, and label all 654 R2V prompts as live-action (486) or animation (168).

Figure 12: Left: T2V single-shot (hatched) vs multi-shot (solid) mean scores (0–100) per model, ordered by multi-shot score; red numbers give the drop ; /. Right: R2V live-action (hatched) vs animation (solid) mean scores per model; /.

Multi-shot is uniformly harder. Every model scores lower on multi-shot prompts (Figure 12, left), with an average drop of 7.9 points (89.2 to 81.3). The drop is larger for lower-ranked models: the top models lose little (Seedance 2.0 , HappyHorse 1.1 ), whereas lower-ranked models show much larger drops (Grok Imagine Video , Veo 3.1 , Hailuo 2.3 ). The degradation concentrates on shot-craft dimensions (Table 5 in the appendix): Cinematic Language (composition, viewing angle, camera movement, focus), editing fluency, and cross-shot spatial consistency, while audio and temporal-coherence dimensions move little. A single-shot prompt only tests intra-shot rendering, whereas a multi-shot prompt additionally requires planning distinct shots and cutting between them inside one clip; the wider inter-model spread on multi-shot prompts reflects a compositional competence that current generators have not yet mastered, and we therefore treat shot structure as a first-class evaluation axis rather than folding it into a single leaderboard number.

Action scenes uniformly lower scores across all models. Every model scores lower on action than on dialogue prompts (Figure 13), and, as with multi-shot prompts, the drop is larger for lower-ranked models: the top group loses least (T2V Seedance 2.0 , HappyHorse 1.1 ) while lower-ranked models show larger drops (Vidu Q3 Pro , Hailuo 2.3 ). Seedance 2.0 maintains the lead on action-specific L3 sub-metrics: on T2V action prompts it tops 16/35 sub-metrics, notably character action (99.3), character expression (96.7), fore/mid/background (95.0) and viewing angle (91.1); on R2V action prompts it tops 11/38 sub-metrics, notably tone & color (100.0), camera movement (84.9), composition (83.1) and shot scale (78.2). Its action drops on these dimensions remain moderate (T2V 3.6, R2V 10.7), indicating that its Cinematic Language and character-performance strengths carry over to action content. The action performance degradation concentrates in Instruction Following and Aesthetic Quality, with two distinct L3 clusters (Table 6 in the appendix). In Instruction Following the drop is dominated by camera-work sub-metrics—camera movement (), focus (), shot scale () and composition ()—indicating that models show larger drops in camera framing and motion once the scene becomes fast and multi-agent. In Aesthetic Quality a second cluster of physical-realism sub-metrics degrades sharply: physical plausibility (), character-motion realism (), perspective (), artifacts () and clarity (). Action scenes thus reveal two concurrent degradation patterns that an aggregate score hides—models not only shift camera framing, but also show reduced physical plausibility and increased artifacts under vigorous motion. The same two clusters lead the R2V action drop, at a smaller magnitude (physical plausibility , artifacts ); we report this contrast descriptively, as the two tasks differ in prompt and sample distribution.

Top models already generalize across visual style. On R2V, the leading group scores almost identically on live-action and animation (Figure 12, right; Seedance 2.0 vs. , gap ), indicating that rendering style is no longer a constraint at the frontier. Style generalization instead shows wider variation in the mid-tier: several mid-ranked models score markedly lower on animation than live-action (Kling 3.0 Omni , ), and the loss again concentrates in Instruction Following (Kling 3.0 Omni ) and Aesthetic Quality, echoing the action analysis: the two harder distributions, action and animation, both challenge models on the same two axes.

Figure 13: T2V (left, 9 models) and R2V (right, 7 models): dialogue (hatched) vs. action (solid) mean scores (0–100) per model, ordered by dialogue score; red numbers give the drop (action dialogue). T2V: action , dialogue . R2V: action , dialogue .

5.10 R2V: Reference Types

R2V prompts condition on one of three reference types: scene (232 prompts), prop (209) and character (213). Figure 14 reports the reference-fidelity (visual-following) ranking for each type. The visual-following leader differs from the overall leader: HappyHorse 1.0 tops scene- and character-fidelity, while HappyHorse 1.1 leads on prop-fidelity, indicating that raw reference fidelity and overall film-quality are distinct capabilities. Across the three types, scene-space and character-appearance fidelity show a similar ranking order, whereas prop-reference fidelity reveals a notably different model ordering and a much tighter spread, confirming that current models still struggle to faithfully preserve a referenced prop even when scene and character references are respected.

Figure 14: R2V visual-following (reference-fidelity) ranking by reference type (columns: scene / prop / character).

5.11 Cinematic Language Related L3 Landscape

Finally, we zoom in on the L3 sub-metrics most directly tied to film production craft. Figure 15 collects 10 sub-metrics from three L2 groups across the two tasks: Cinematic Language (camera movement, shot scale, focus, viewing angle, tone & color, composition), editing (camera-work appeal, editing fluency), and performance (action performance, emotional performance). Across both tasks, models score highest on tone & color (T2V 95.0, R2V 91.2) and lowest on action performance (50.8 / 46.8) and camera-work appeal (54.8 / 57.5). The performance group is uniformly the lowest-scoring L2 cluster in both tasks, while Cinematic Language dimensions show the widest T2V–R2V gap (e.g., camera movement 69.5 vs. 55.6), indicating that reference conditioning amplifies the spread on camera-work execution.

Figure 15: Cinematic Language L3 sub-metric cross-model means (0–100), T2V (hatched) vs. R2V (solid), sorted by descending score. Bars are colored by L2 group: Cinematic Language (blue), editing (red), performance (purple).

5.12 Summary of Findings

Taken together, the analyses above distill into five findings that hold across the field. (i) No saturation; models differ most on Cinematic Language sub-metrics: no model nears the ceiling on either task (tops of 88.93 / 86.66), and camera movement, focus, shot scale and viewing angle show the largest inter-model spread, far more so under R2V (camera-movement variance vs. ), while temporal and audio dimensions show near-uniform scores. (ii) Dynamic aesthetics are the lowest-scoring sub-metrics across the field: action performance, camera-work a

正文较长,站内已展示前半部分,完整内容请阅读原文。
Hugging Face多模态开源/仓库视频
阅读原文导出 Markdown

Hengxia Qiang

Weibin Chen

Jinyang Zhen

Bing Zhao

Jing Li

Hu Wei

Moku Lab, Hujing Digital Media & Entertainment Group

Beijing Film Academy

qihong@bfa.edu.cn

kongwang@alibaba-inc.com

lj225205@alibaba-inc.com

Abstract

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film-academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman (T2V) and (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.

1 Introduction

Figure 1: FilmBench evaluation taxonomy: 3 L1 axes, 12 L2 components, 35+3 (R2V-only) L3 sub-metrics, built from clips across 20 film genres curated by expert directors.

Figure 2: The director-driven reverse-engineering pipeline that turns award-winning film clips into film-grade prompts; see §3.1 for details.

Video generation has progressed at an unprecedented pace. Diffusion- and DiT-based models (Veo [4, 24], Kling [12], Seedance [20], Hailuo [17], HappyHorse [6, 7] and others) now produce 15-second photorealistic clips with controllable camera and even synchronized audio, blurring the line between AI-generated footage and professionally produced cinema. A natural next question is whether the capability of these models has reached the pass mark of professional cinematic creation. Modern film production follows a rigorous Cinematic Language (shot scale, camera movement, shot perspective, lighting, scene blocking, composition and film-grade audio design) refined over more than a century of practice. Existing benchmarks were not designed with this professional, academy-taught Cinematic Language in mind, and so far they paint only a coarse, optimistic picture of model capability.

The current benchmarking landscape, and what is missing.

The community has produced a rich body of evaluation suites in the past three years. Foundational multi-dimensional benchmarks such as VBench [9], VBench-2.0 [29] and Video-Bench [5] dissect quality into 9–18 generic dimensions, and learned scorers such as VideoScore [8] push automatic evaluation closer to human preference. Compositional and fine-grained benchmarks (T2V-CompBench [22], FETV [15]) target attribute, motion and interaction binding; temporal, motion and multi-shot benchmarks (ChronoMagic-Bench [27], VMBench [14], SLVMEval [16], MSVBench [21]) probe long-horizon, motion and cross-shot fidelity; reasoning, physical and social benchmarks (TiViBench [2], SVBench [18], WorldJen [10], RBench [3]) stress higher-order or embodied capabilities; the I2V line (ConsistI2V/I2V-Bench [19], UI2V-Bench [28], IP-Bench [13]) focuses on image-conditioned generation; and audio–video and aesthetic benchmarks (AVGen-Bench [30], VGA-Bench [11]) address audio quality and visual aesthetics, respectively.

Figure 3: Fine-grained, per-dimension FilmBench evaluation on a multi-shot reference-scene (R2V) example reverse-engineered from La La Land: Seedance 2.0 (94.97) vs. Vidu Q2 Pro (73.61). Discussed in the text.

Despite this breadth, none of these benchmarks evaluates models against the standards used in real cinematic creation. They share three structural blind spots. First, prompts are sourced from web users, captioners or LLM templates rather than from verified professional shots, so the input distribution drifts away from cinematic content. Second, the implicit content distribution skews toward generic, English, web-style scenes, with no balanced coverage of cinematic genres or culturally diverse film traditions. Third, the evaluation taxonomies leave out cinematic axes such as shot scale, camera movement, lighting, visual style, character performance and audio quality, or collapse them into a single “aesthetic” score. As a consequence, current leaderboards saturate while professional cinematic creators readily reject the same outputs as not film-grade.

FilmBench.

We address this gap with FilmBench, an evaluation benchmark for cinematic text-to-video (T2V) and reference-to-video (R2V) generation, built jointly with directors and faculty from the Beijing Film Academy and the film studio of Hujing Digital Media & Entertainment Group. FilmBench is grounded in three design principles.

Figure 1 gives an overview of the FilmBench evaluation taxonomy. Expert directors from the Beijing Film Academy selected clips spanning 20 film categories—covering the breadth of professional cinematic genres—and used them as the foundation for both the prompt set and the fine-grained evaluation hierarchy. The taxonomy radiates from the three L1 axes through progressively finer L2 components and L3 sub-metrics, while the surrounding film examples illustrate the kinds of cinematic-language challenges each dimension targets. We next describe how these expert-curated clips are turned into structured, film-grade prompts.

(P1) Reverse-engineered from real films (Figure 2): instead of authoring prompts in the abstract, professional directors select clips from award-winning films that stress Cinematic Language and visual expression, reverse-infer draft prompts with a strong multimodal model (Gemini 3.1 Pro), and refine them into structured shot-level prompts that explicitly encode scene, role, prop, shot scale, camera movement, composition, lighting, dialogue and performance, so that every prompt is anchored to a verified professional reference and every fine-grained criterion in the taxonomy is covered. Because these prompts follow real shot lists, most are multi-shot (1,056 of the 1,169 prompts), unlike the single-clip prompts of prior benchmarks. (P2) Academy-aligned cinematic taxonomy: our evaluation dimensions follow the Cinematic Language teaching system of the Beijing Film Academy, organized as three first-level axes (L1): Instruction Following, Temporal Continuity and Aesthetic Quality. These decompose into 12 second-level components (L2) and 35 third-level sub-metrics (L3). T2V and R2V share this main framework; the R2V task additionally adds a Visual Following L2 component of three L3 sub-metrics (scene-space, character-appearance and prop-reference fidelity), each evaluated against the conditioning reference image, giving R2V 13 L2 and 35+3 L3 (per prompt, only the sub-metric matching its reference type is scored). (P3) Expert-grade automatic evaluator: every sub-metric is scored by an in-house, professional film-grade evaluation agent whose core Cinematic Language operator suite (FilmOps) we open-source; its model-level ranking reproduces expert film-industry rankings at Spearman (T2V) / (R2V) (§3, §4).

A worked example.

Figure 3 shows FilmBench scoring a multi-shot reference-scene (R2V) example, a four-shot clip reverse-engineered from the musical La La Land, on Seedance 2.0 (94.97) and Vidu Q2 Pro (73.61). Because every L3 sub-metric is scored independently, the evaluation exposes behavior that an aggregate score hides. Although Seedance 2.0 leads overall by a wide margin, its scene-space fidelity is slightly lower (87.5 vs. 100): Vidu Q2 Pro over-adheres to the reference background image and scores lower on nearly every other dimension. It drops markedly on Cinematic Language following (shot scale 75 vs. 25, viewing angle 100 vs. 50, camera movement 100 vs. 50, composition 75 vs. 43.75) and, having sacrificed the shot/reverse-shot staging to lock onto the scene, exhibits character-positioning drift that drives its character blocking to 0. FilmBench can thus objectively credit a single dimension (reference fidelity) while still penalizing the broader loss of professional Cinematic Language; Appendix A discusses further per-dimension case studies.

Findings.

We benchmark leading video-generation models (9 for T2V, 7 for R2V), including Seedance 2.0, HappyHorse 1.1/1.0, Kling 3.0, Kling 3.0 Omni, Veo 3.1, Grok Imagine Video [26], Vidu [1], and Hailuo 2.3, and distill five field-wide findings. (i) No saturation: no model approaches the ceiling on either task (top scores of 88.93 for T2V and 86.66 for R2V), in sharp contrast with the near-ceiling numbers on prior web-style benchmarks; and crucially, what separates the leaders is not generic quality but the professional Cinematic Language sub-metrics of Instruction Following (camera movement, shot scale, viewing angle, focus and composition) together with the Aesthetic Quality axis, which is exactly where FilmBench resolves differences that web-style benchmarks miss. (ii) A field-wide dynamic-aesthetics bottleneck: across all models the lowest-scoring sub-metrics are action performance, camera-work appeal, character-motion realism and emotional performance, whereas static image-quality sub-metrics (sharpness, lighting) score markedly higher: today’s models render clean frames but still struggle with believable motion and performance. (iii) Multi-shot is the hardest regime: on T2V, moving from single- to multi-shot prompts costs points on average and up to for the weakest model, concentrated in the Cinematic Language and editing demands of Instruction Following (framing, camera work and cut fluency); it is multi-shot staging, rather than single-take quality, where current models fall short. (iv) Reference conditioning stresses rather than reorders the field: across scene, prop and character references the R2V ranking stays stable, yet conditioning lowers scores to varying degrees and penalizes the weakest model most, and the visual-following leader is not necessarily the overall leader, so reference behavior is best read through ranking shifts rather than absolute per-type scores. (v) No model wins on all L3 sub-metrics: Seedance 2.0, the overall leader, claims only 18/35 (T2V) or 20/38 (R2V) L3 championships, concentrated on Cinematic Language sub-metrics (Cinematic Language, editing appeal and performance), while HappyHorse 1.1 dominates 11/35 (T2V) or 10/38 (R2V) sub-metrics on character, audio and scene dimensions; notably the three R2V-specific visual-following championships all go to the HappyHorse family; 4 (T2V) and 2 (R2V) models do not claim any championship. This “championship mismatch” reveals structural complementarity concealed by a single leaderboard number.

Contributions.

We make three contributions. (i) We release FilmBench, a cinematic-grade T2V/R2V benchmark whose prompts are reverse-engineered from award-winning real films and curated jointly with directors and faculty from the Beijing Film Academy and a professional film studio, ensuring coverage of every fine-grained cinematic criterion. (ii) We release FilmOps, a modular Cinematic Language operator suite with trained weights and inference scripts, enabling the community to build customizable, expert-knowledge-driven judge agents for video evaluation. (iii) We propose an academy-aligned evaluation taxonomy that translates professional Cinematic Language into a three-level hierarchy of 3 L1 axes, 12 L2 components and 35 L3 sub-metrics (35+3 with the R2V-specific Visual-Following extension), scored by an expert-grade automatic evaluator with an open-source Cinematic Language operator suite. (iv) We conduct a comprehensive cinematic evaluation over leading T2V/R2V models, show the automatic evaluator reproduces expert rankings at (T2V) / (R2V), and expose dynamic aesthetics as the shared bottleneck, discussing directions for future research.

2 Related Work

We trace how video-generation evaluation has evolved through three stages (from generic quality scoring, to capability-specific stress tests, to increasingly film-like assessment) and show that each stage, in closing one gap, exposes the next, until the accumulated gaps motivate FilmBench precisely.

From distribution metrics to multi-dimensional, human-aligned evaluation.

Early evaluation repurposed action-recognition datasets (UCF-101, MSR-VTT, Kinetics) and reported single scalars such as FVD/IS, which neither localize which aspect of a video fails nor align well with human perception. As diffusion- and DiT-based generators—including the systems we evaluate (Veo-3.1, Kling-v3/Omni, Seedance-2.0, Hailuo-2.3, Vidu-Q3-Pro, Grok-Imagine-Video and HappyHorse-1.x)—raised the ceiling to controllable, audio-equipped clips, the community moved to disentangled, human-aligned protocols. VBench [9] and VBench-2.0 [29] decompose quality into 16 then 18 dimensions and push from superficial to intrinsic faithfulness; Video-Bench [5] enlists MLLMs as scalable judges; VideoScore [8] learns a human-aligned regressor. This stage established the grammar of modern evaluation—quality as a structured set of interpretable dimensions—but its axes are deliberately generic and its prompts are drawn from web users, reflecting what is easy to source rather than what a film production demands. The first gap is thus one of content and taxonomy: neither the prompts nor the dimensions are anchored in a professional production grammar.

Specializing to emerging capabilities.

With this foundation in place and baseline visual quality saturating, benchmarks specialized to isolate the hard capabilities that generic scores had masked. Compositional suites T2V-CompBench [22] and FETV [15] expose brittle attribute, motion and spatial-relation binding; the temporal–motion line ChronoMagic-Bench [27] and VMBench [14] target metamorphic amplitude and perception-aligned motion, while SLVMEval [16] extends reliability testing to hour-scale footage. As generators began stitching shots into a story, MSVBench [21] introduced multi-shot evaluation with a hybrid LMM-plus-expert scorer, distilling this stage’s lesson: current systems behave as visual interpolators rather than world models. A parallel reasoning wave then raised the bar from rendering to understanding, probing structural (TiViBench [2]), social (SVBench [18]), broad multi-axis (WorldJen [10]) and embodied-robotic (RBench [3]) reasoning. Yet even MSVBench scores its shots against generic quality criteria rather than a director’s intended shot list, so evaluation still stops short of the film: cross-shot narrative structure and director intent are not judged end-to-end against a verified cinematic reference. This is the second gap.

Toward film: reference, audio, aesthetics and Cinematic Language.

The strand closest to our goal relaxes that assumption toward controllable, film-like generation, but does so one aspect at a time. Reference/image-conditioned suites (ConsistI2V/I2V-Bench [19], UI2V-Bench [28] and IP-Bench [13]) evaluate consistency, semantic understanding and protection from a conditioning image, raising reference fidelity as a concern. As generators acquired joint audio, AVGen-Bench [30] evaluates text-to-audio-video generation, adding a cross-modal audio axis, while VGA-Bench [11], co-developed with the Beijing Film Academy, brings aesthetic tagging into view. In parallel, a cinematic-understanding literature (MovieNet, shot-type taxonomies, CineScale/ShotBench) and storyboard-to-video pipelines treat film as structured language, and MovieBench [25] pushes to feature-length identity consistency—sustained in at most of cases at multi-scene scope. Collectively these efforts touch nearly every ingredient of film, yet no single benchmark unifies them: aesthetics and audio are not anchored in a production grammar, reference fidelity and cross-shot audio–dialogue continuity are evaluated in isolation if at all, the target is typically real or arbitrary footage rather than generated films, and the cinematic-understanding line asks only whether models can read Cinematic Language, never whether a generator can speak it at a director’s granularity. Closing this third gap—a unified, production-grounded evaluation of generated film—is exactly the remit of FilmBench.

Positioning of FilmBench.

These three accumulated gaps (a generic taxonomy, a clip-level scope, and fragmented coverage of individual film aspects) jointly define what a film-grade benchmark must satisfy at once. To the best of our knowledge, no existing benchmark simultaneously (i) reverse-engineers prompts from genuine, award-winning films selected by directors; (ii) spans a broad set of cinematic genres and both text-to-video and reference-to-video tasks, with prompts organized as real, mostly multi-shot shot lists (1,056 of 1,169) that evaluate the film rather than the isolated clip; (iii) grounds its taxonomy in the Cinematic Language system taught in professional film schools: 3 first-level axes, 12 second-level components and 35 third-level sub-metrics with an R2V-specific Visual-Following extension, co-designed with directors, Beijing Film Academy faculty and a professional film studio; and (iv) treats audio–dialogue continuity (including dialogue intelligibility and per-character timbre stability across shots) as a first-class axis. FilmBench thus asks not whether a model can read Cinematic Language but whether it can speak it, and because every prompt is reverse-engineered from a verified shot, we always retain both a structured grammar and a ground-truth reference. §3 and §4 detail this taxonomy and the automatic-plus-human evaluator stack, with a head-to-head comparison table in §3 and the appendix.

3 FilmBench: Construction and Dimensions

This section describes how FilmBench is built. We first present the reverse-engineering pipeline (Figure 2) that turns genuine cinematic clips into structured prompts (§3.1), then introduce our three-level Cinematic Language taxonomy of 3 first-level axes (L1), 12 second-level components (L2) and 35 third-level sub-metrics (L3), co-designed with faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio (§3.2).

3.1 Reverse-Engineered Prompt Construction

FilmBench prompts are produced by a director-driven reverse-engineering pipeline with three stages (Figure 2). Clip selection: professional directors from the film studio of Hujing and the Beijing Film Academy prioritize award-winning films (e.g., winners of the Academy Awards, the Golden Horse Awards and the Hundred Flowers Awards) across our 20 cinematic genres, and select multi-shot clips (mostly shots) that specifically stress Cinematic Language and visual-expression skill. Draft prompt inference: a FilmOps operator suite together with a strong multimodal model (Gemini 3.1 Pro) reverse-extracts the original narrative script, Cinematic Language elements and on-screen visual elements from each clip, which Gemini 3.1 Pro then assembles into an initial structured prompt. Expert refinement: the Beijing Film Academy directing team screens, quality-checks, rewrites and standardizes every draft into professional film-grade Cinematic Language. Throughout clip selection and final curation, the team ensures that every L1 axis and L2 component in the taxonomy is covered, so that the benchmark faithfully probes a model’s film-grade generation ability rather than generic visual quality. Each prompt uses slotted cinematic tags (@scene, @role, @prop), e.g., “Shot 1 (medium-shot, eye-level fixed, rule-of-thirds): in a brightly lit ICU @scene1, a young female doctor @role1 ”.

Tasks and scale.

FilmBench covers two tasks: T2V (prompt video) and R2V (reference prompt video). The T2V suite contains 515 prompts spanning 20 base cinematic genres, balanced across Chinese Films and International Films markets (243 Chinese Films / 272 International Films). The R2V suite contains 654 prompts split across three reference types (232 scenes / 209 props / 213 characters) and drawn from the same 20-genre pool across both markets (239 Chinese Films / 415 International Films). Because prompts follow real shot lists, most are multi-shot: 402 of the 515 T2V prompts and all 654 R2V prompts script multiple shots (1,056 of 1,169 overall), unlike the single-clip prompts of prior benchmarks. Each prompt is reverse-engineered from a distinct award-winning film clip, so the prompt count also reflects the number of source clips.

3.2 Academy-Aligned Evaluation Taxonomy

Our evaluation dimensions follow the academic Cinematic Language system taught in professional film schools. The taxonomy is co-designed by faculty from the Beijing Film Academy and the film studio of Hujing together with an analysis of current video-generation models’ capabilities, and is deliberately organized as a multi-level hierarchy so that conclusions and insights can be drawn at coarse (L1), medium (L2) and fine (L3) granularity.

媒体内容 · 前往原文查看

Table 1: FilmBench three-level taxonomy (3 L1 axes, 12 L2 components, 35 L3 sub-metrics). R2V adds a Visual Following component (+3 L3). Full definitions in Appendix B.

L1 Axis L2 Component

Instruction Following (16+3) Cinematic Language

Character & performance

Scene

Visual Following (R2V only)

Temporal Continuity (7) Spatial coherence

Temporal coherence

Character coherence

Audio coherence

Aesthetic Quality (12) Base quality

Editing

Performance

Audio quality

Three L1 axes.

FilmBench evaluates cinematic competence along three first-level axes, decomposed into 12 L2 components and 35+3 L3 sub-metrics (Table 1); below we enumerate the Cinematic Language sub-metrics and leave the remaining, more generic components to Table 1 and Appendix B. Instruction Following captures whether the video faithfully realizes the prompt’s cinematic intent, over four components: Cinematic Language (shot scale, camera movement, viewing angle, composition, focus, tone & color, thematic style), Character & performance, Scene, and Audio. Temporal Continuity captures whether the output behaves as a film rather than a single clip, over four components: spatial coherence, temporal coherence, character coherence, and audio coherence. Aesthetic Quality captures production-grade quality, over four components: base quality, editing (editing fluency, camera-work appeal), performance (emotional performance, action performance), and audio quality.

Task-specific extension.

T2V and R2V share this main framework. R2V additionally introduces one L2 component, Visual Following, with three L3 sub-metrics (scene-space, character-appearance and prop-reference fidelity), each measuring adherence to the conditioning reference image. Every R2V prompt conditions on one of three reference types (scene, prop or character), and only the matching fidelity sub-metric is scored for that prompt. Visual Following is scored only for R2V, so R2V has 13 L2 and 35+3 L3 dimensions, while all other dimensions are identical across tasks. For every L3 sub-metric we provide a 5-point anchor rubric; the complete definition table is in Appendix B.

4 Evaluation Method

FilmBench is scored by an expert-grade automatic evaluation agent which integrates a suite of Cinematic Language operators, expert annotation models, and a judge model. For every dimension in the taxonomy the agent produces a 1–5 score, achieving high agreement with human experts (§4.2). Below we describe the open-sourced operator suite (§4.1) that grounds the Cinematic Language dimensions and the score-aggregation rule (§4.2).

4.1 FilmOps: A Cinematic Language Operator Suite

A key obstacle to professional cinematic evaluation is that neither generic MLLM-as-judge approaches nor existing domain expert models reliably encode film-industry visual rules: general multimodal models are trained on open-domain understanding and misjudge professional categories such as shot scale, composition and camera movement [23], while existing aesthetic/shot models are trained mostly on everyday photography or live-action footage and do not transfer across genres (3D/2D animation, stylized content). We therefore build FilmOps, an open-source operator suite that maps a generated video into structured cinematic labels for the Cinematic Language dimensions.

Industry-aligned taxonomy.

FilmOps defines a professional classification system, aligned with classical production references and vetted by front-line practitioners, over six core dimensions: shot scale, composition, viewing angle, tone & color, character layout, and camera movement. Five are frame-level and one (camera movement) is shot-level; together they span 55 sub-categories (character layout is an open-ended natural-language field). Categories are defined by narrative/visual function rather than imaging mechanism, so the same standard applies across live-action, 3D and 2D genres.

Multi-genre training data.

To match the cross-genre, multi-style distribution of generated video, FilmOps is trained on a pool of 5,000+ real film/TV works spanning live-action, 3D animation, 2D animation and stylized/VFX content, and across narrative types (dialogue-driven, action, group scenes, and long-tail shot forms). Each operator is trained on 40K–60K annotated samples with a shot-based test set, under a strict annotator-training and quality-control protocol led by professional practitioners.

Task-matched operator design.

Operators are built to match each dimension’s characteristics: frame-level visual dimensions (shot scale, composition, angle, color) use vision backbones (DINO, BEiT and ResNet-18), while dimensions requiring relational reasoning or temporal modeling (character layout, camera movement) use a multimodal model (InternVL3-14B, LoRA-SFT). Classification operators are evaluated by macro-F1 and the natural-language layout operator by precision. Against four strong zero-shot general-MLLM baselines, the trained operators improve markedly on every dimension, confirming the value of domain-specialized operators; the full per-operator backbones and macro-F1 figures tested over 400 images/shots against these baselines are reported in the appendix (Table 12).

To support customizable, personalized self-evaluation, we open-source the classification standard, model weights and inference scripts of FilmOps. The modular design is not tied to a single model, so users can reuse individual operators or embed specific dimensions into their own evaluation pipelines.

4.2 Score Aggregation

Each L3 sub-metric receives a 1–5 score, which is linearly mapped to a 0–100 scale before aggregation (). Aggregation is sample-level: each video is first aggregated across its dimensions, and per-model scores are the mean over that model’s videos, preserving sample variance for significance analysis. We aggregate L3 L1 directly (each L1 axis is the mean of its constituent L3 sub-metrics), treating every L3 sub-metric as equally important, and the Overall score is the equal-weighted mean of the three L1 axes; L2 components are computed as a side branch for fine-grained reporting only. Not-applicable scores (N/A or null) are excluded from aggregation; for models without audio capability, audio dimensions are handled symmetrically as not-applicable.

Qualitative case studies.

Fine-grained per-dimension scoring exposes trade-offs that aggregate scores hide. As a worked example, Figure 4 scores a multi-shot T2V sci-fi mech battle, a case that stresses the dual challenge of intense action combined with multi-shot cutting. Seedance 2.0 (86.11) clearly outperforms Grok Imagine Video (55.97), and the gap is diagnostic rather than uniform: Seedance is markedly stronger on tone & color, character & performance and scene, and on camera-work appeal; in the action passages its character positioning stays consistent across shots and its temporal logic shows no visible errors. Grok, by contrast, exhibits scene drift and drifting cross-shot appearance and positioning, which surfaces directly in the per-dimension gaps on character-motion realism and action performance. The introduction (Figure 3) shows a complementary multi-shot R2V reference-scene case, and Appendix A provides two further worked examples (reference-character and reference-prop) that together cover the major scoring scenarios in FilmBench.

Figure 4: Multi-shot T2V example (a sci-fi mech battle): Seedance 2.0 (86.11) vs. Grok Imagine Video (55.97). Discussed in the text.

5 Experiments

5.1 Models and Setup

We benchmark leading video-generation models under each task. For T2V we evaluate 9 models: Seedance 2.0 [20], HappyHorse 1.1 [7], HappyHorse 1.0 [6], Kling 3.0 [12], Kling 3.0 Omni [12], Veo 3.1 [4, 24], Grok Imagine Video [26], Vidu Q3 Pro [1] and Hailuo 2.3 [17]. For R2V we evaluate 7 models: Seedance 2.0, HappyHorse 1.1, HappyHorse 1.0, Kling 3.0 Omni, Veo 3.1, Grok Imagine Video, and Vidu Q2 Pro [1]. All models are run under their official APIs / public checkpoints with default settings, generating one video per prompt. This yields machine-scored T2V records and R2V records. Throughout, T2V and R2V are reported with the same analysis format: overall, per-axis (L1), per-component (L2), per-sub-metric (L3), discriminability, and capability profiles, so that the two tasks are directly comparable; R2V additionally carries a reference-type analysis (scene / character / prop).

5.2 Human Agreement

We validate the automatic evaluator against 90 professional human raters majoring in film, television, directing, or digital media from the Beijing Film Academy at the model level: over 300 prompts (30% of the benchmark) carry both machine and human judgments, yielding model–prompt pairs. We compute per-model machine and human means at each L2 component and correlate the two rankings (Table 2). The overall model-level Spearman reaches (T2V) and (R2V), confirming that the automatic leaderboard closely reproduces expert model rankings.

媒体内容 · 前往原文查看

Table 2: L2 component-level human agreement (Spearman ), averaged across T2V and R2V. 10 of 13 components reach ; the three lower-agreement components (audio quality, audio coherence, editing) are separated by a mid-rule.

Perform.

Vis. Follow.

Cinem. Lang.

Base qual.

Char. & perf.

Temp. coh.

Spat. coh.

Scene

Char. coh.

Audio q.

Audio coh.

Editing

0.99 0.96 0.95 0.95 0.93 0.92 0.85 0.83 0.82 0.78 0.68 0.68 0.60

Drilling into the component level, 10 of the 13 L2 dimensions achieve ; the three lower-agreement components—audio quality (), audio coherence () and editing ()—are those where subjective temporal behaviour (sound fidelity, cut rhythm) is hardest to pin down, and we acknowledge that human and automatic judgments diverge more on these axes. The overall-level agreement ( on both tasks) nonetheless indicates that the aggregate ranking is a reliable proxy for expert judgment, and we flag the three lower-agreement dimensions when interpreting fine-grained results.

5.3 Overall Leaderboard

Tables 7–8 and Figure 5 report the overall FilmBench score (0–100, higher is better). On T2V, Seedance 2.0 leads (88.93), closely followed by HappyHorse 1.1 (87.42) and HappyHorse 1.0 (87.02); the two Kling 3.0 variants form a second tier (86), Vidu Q3 Pro, Grok and Veo 3.1 a third tier (81), with Hailuo 2.3 at 68.94. On R2V, the top group is Seedance 2.0 (86.66), HappyHorse 1.1 (85.51) and HappyHorse 1.0 (84.87). Crucially, no model approaches saturation, in sharp contrast with the near-ceiling scores reported on prior web-style benchmarks, and the top group is consistent across both tasks, indicating that reference conditioning does not reshuffle the leaders.

Figure 5: Overall FilmBench scores (0–100) per model; model names are printed inside each bar and the dashed line is the cross-model average. Left: T2V (); right: R2V ().

5.4 Dimension Variance

To quantify where models differ most at fine-grained levels, we compute the cross-model variance of each dimension across the hierarchy (Figure 6; Hailuo is excluded from the variance computation to avoid its audio outliers dominating). At the axis level, Instruction Following shows by far the largest variance (), nearly eight times that of Aesthetic Quality () and seventeen times that of Temporal Continuity (), indicating that models differ primarily in their ability to execute professional cinematic language instructions.

Drilling into the component level, Cinematic Language dominates with a variance of , far exceeding Audio (), Scene () and all other components. At the sub-metric level, the top five highest-variance dimensions are all Cinematic Language or scene-related: camera movement (), focus (), shot scale (), viewing angle () and fore/mid/background (). This concentration confirms that the primary differentiator among current generative video models lies in Cinematic Language instruction following—whether a model understands and executes professional camera-language conventions. Per-task breakdowns are provided in Appendix C.

Figure 6: Cross-model variance per L3 sub-metric (main bars, colored by axis), with L2 inset (upper right), computed over the merged T2V+R2V model means (Hailuo excluded).

5.5 Per-Axis Landscape (L1)

Across the three axes, models are most differentiated on Instruction Following, which encompasses the Cinematic Language sub-metrics that drive the high variance reported above. On T2V (Figure 7, left), Seedance 2.0 leads all three axes: Instruction Following (91.69, ahead of HappyHorse 1.1 at 90.70), Temporal Continuity (96.06) and Aesthetic Quality (79.03). On R2V (Figure 7, right), Seedance 2.0 likewise leads every axis (Instruction Following 87.45, Temporal Continuity 94.09, Aesthetic Quality 78.44). The Instruction Following axis shows the widest inter-model gap under both tasks, as the Cinematic Language and scene-related sub-metrics amplify the spread in professional-language execution.

Figure 7: Per-axis rankings for T2V (left, 9 models) and R2V (right, 7 models). Each model is scored on three L1 axes: Instruction Following (solid), Temporal Continuity (hatched), and Aesthetic Quality (dotted). Models are ranked by overall score; R2V Instruction Following includes the visual-following component.

5.6 Per-Component Rankings (L2)

Descending one level, we rank models within the L2 components most closely tied to Cinematic Language together with the adjacent Scene and Character & performance components (Figure 8); the full 12/13-panel rankings are deferred to the appendix. Cinematic Language is placed at the centre and used as the sort key, and each component isolates a different craft skill, revealing where the head models specialize:

Cinematic Language (camera work, framing, transitions) is the sharpest discriminator of the three, with the field fanning out over a 17–42-point range and even the strongest model reaching only 84.6. This is the one component that hinges on temporal decision-making—when to move the camera, how to reframe, when to cut—rather than per-frame rendering, which is why models diverge so widely here. Seedance 2.0 leads decisively and is the only model whose camera competence holds up when a reference image is imposed, marking dynamic camera control as its signature strength.

Figure 8: Per-model rankings within the L2 components most related to audiovisual language: Cinematic Language (solid, centre, model names inside; also the sort key), flanked by Scene (hatched) and Character & performance (dotted). T2V (left, 9 models) and R2V (right, 7 models). The full L2 rankings are in the appendix.

Scene (environment consistency, lighting, background layout) also spreads the field, though the leading cluster reaches the mid-90s (HappyHorse 1.0 at 96.0, HappyHorse 1.1 at 96.7). The gap opens further down the ranking, where weaker models struggle to keep a coherent, consistently lit environment across shots. The separation here is owned by the HappyHorse family, whose strength lies specifically in preserving spatial layout—a capability distinct from camera craft, since the models that lead scene rendering are not the ones that lead Cinematic Language.

Character & performance spreads the field the least of the three but still meaningfully, with the leaders reaching the high-90s (HappyHorse 1.1 at 97.6, Seedance 2.0 at 96.3). Here differentiation comes from sustaining believable characters and acting across a full scene, and HappyHorse 1.1 holds a slim but consistent edge, making it the character specialist just as Seedance 2.0 is the camera specialist. Read together, the three components show that the overall leaderboard masks a division of labor—camera craft, scene rendering and character performance are led by different models—so no single system dominates every axis of film craft.

5.7 Fine-Grained Rankings (L3)

At the finest level, FilmBench resolves 35 (T2V) / 38 (R2V) L3 sub-metrics, each with its own ranking; Figure 9 shows the three sub-metrics selected by highest merged T2V+R2V variance. All three are Cinematic Language sub-metrics, indicating that camera craft is where the leading models pull apart most:

Camera movement shows the largest spread among the head models, with scores ranging from Seedance 2.0’s 86.5 down toward the low 50s. Executing a motivated camera move—a push-in, a pan that tracks action, a crane reveal—requires the model to sustain a coherent scene through time, and models diverge considerably in how convincingly they do so. Seedance 2.0 leads by a clear margin, suggesting a training emphasis on dynamic camera choreography that sets it apart from the rest of the field.

Figure 9: Per-model rankings for the three highest-variance L3 sub-metrics (computed over merged T2V+R2V data): Camera movement (hatched), Focus (solid, model names inside), and Shot scale (dotted). T2V (left, 9 models) and R2V (right, 7 models). The full per-metric panels are in the appendix.

Focus (depth-of-field control, subject isolation) sees the leaders cluster in the low-to-mid 80s (Seedance 2.0 at 85.6, HappyHorse 1.1 at 82.7), with a wider tail below. The metric separates models because it rewards intentional focus that directs attention to the dramatically relevant subject—a semantic judgement, beyond merely producing a shallow-depth look, that the stronger models handle more reliably.

Shot scale (close-up, medium, wide framing decisions) tests whether a model chooses the framing a scene calls for—an intimate close-up for a line of dialogue, a wide for an establishing beat. Seedance 2.0 again leads (79.8) with HappyHorse 1.1 the runner-up, and the head ordering across all three sub-metrics is consistent—Seedance 2.0 first, the HappyHorse family close behind. Because these are the dimensions on which models differ most, holding a high standard across them is what distinguishes a film-grade leader: Seedance 2.0’s overall lead comes not from a single capability but from staying near the top on the most challenging Cinematic Language dimensions at once, while the HappyHorse family remains competitive across them.

Crucially, no model wins on all L3 sub-metrics: Seedance 2.0, the overall leader, claims only 18/35 (T2V) or 20/38 (R2V) championships, concentrated on Cinematic Language sub-metrics—Cinematic Language (camera movement, shot scale, focus, viewing angle), camera-work appeal, and performance (action and emotional performance)—in both tasks. HappyHorse 1.1, the runner-up, dominates 11/35 (T2V) or 10/38 (R2V) sub-metrics on character instruction-following, audio, and scene dimensions. Notably, on the three R2V-specific visual-following sub-metrics (scene-space fidelity, character-appearance fidelity and prop-reference fidelity), the HappyHorse family takes all three championships—HappyHorse 1.0 leads scene-space and character-appearance fidelity, HappyHorse 1.1 leads prop-reference fidelity. This “championship mismatch” is the key value of fine-grained evaluation: a single leaderboard number conceals the structural complementarity between models. Beyond the top two, the remaining championships are scattered across several models, with 4 (T2V) or 2 (R2V) models winning no championships at all. This granularity turns a single leaderboard into an actionable diagnostic: a mid-tier model can still lead on individual craft dimensions. The complete per-model score matrices (L3 and L2 heatmaps) are in the appendix (Figures 20–23).

5.8 Market and Genre Robustness

FilmBench spans 20 cinematic genres across Chinese Films and International Films markets (T2V /; R2V /) to ensure broad diversity and coverage. Figure 10 splits each task by market: the top group (Seedance 2.0, HappyHorse 1.1/1.0) is identical for Chinese Films and International Films prompts on both tasks, with only minor mid-pack reordering. Figure 11 shows the score heatmap across the top 4 genres by sample size: rankings remain stable across genres, and each model’s per-genre spread is small (2–3 points), confirming that both leaderboards reflect model capability rather than genre mix. The action subset, often cited as the hardest, shows the same ranking with only a slightly lower ceiling.

Figure 10: Overall ranking split by market, Chinese Films vs. International Films; left: T2V (/), right: R2V (/).

Figure 11: Overall score heatmap across the top 4 genres by sample size. Y-axis: genres (with sample sizes T2V/R2V); X-axis: models split into T2V (left) and R2V (right) sections. Each cell shows the genre-specific mean score.

5.9 Prompt Complexity and Content Factors

Beyond market and genre, we probe three prompt-level factors that stress generation differently: narrative structure (single- vs multi-shot prompts), content type (action vs dialogue scenes), and rendering style (live-action vs animation). In T2V, 113 of 515 prompts are single-shot and 402 are multi-shot; every R2V prompt is multi-shot, so the shot contrast is T2V only. We tag 117 of 515 T2V prompts and 190 of 654 R2V prompts as action, and label all 654 R2V prompts as live-action (486) or animation (168).

Figure 12: Left: T2V single-shot (hatched) vs multi-shot (solid) mean scores (0–100) per model, ordered by multi-shot score; red numbers give the drop ; /. Right: R2V live-action (hatched) vs animation (solid) mean scores per model; /.

Multi-shot is uniformly harder. Every model scores lower on multi-shot prompts (Figure 12, left), with an average drop of 7.9 points (89.2 to 81.3). The drop is larger for lower-ranked models: the top models lose little (Seedance 2.0 , HappyHorse 1.1 ), whereas lower-ranked models show much larger drops (Grok Imagine Video , Veo 3.1 , Hailuo 2.3 ). The degradation concentrates on shot-craft dimensions (Table 5 in the appendix): Cinematic Language (composition, viewing angle, camera movement, focus), editing fluency, and cross-shot spatial consistency, while audio and temporal-coherence dimensions move little. A single-shot prompt only tests intra-shot rendering, whereas a multi-shot prompt additionally requires planning distinct shots and cutting between them inside one clip; the wider inter-model spread on multi-shot prompts reflects a compositional competence that current generators have not yet mastered, and we therefore treat shot structure as a first-class evaluation axis rather than folding it into a single leaderboard number.

Action scenes uniformly lower scores across all models. Every model scores lower on action than on dialogue prompts (Figure 13), and, as with multi-shot prompts, the drop is larger for lower-ranked models: the top group loses least (T2V Seedance 2.0 , HappyHorse 1.1 ) while lower-ranked models show larger drops (Vidu Q3 Pro , Hailuo 2.3 ). Seedance 2.0 maintains the lead on action-specific L3 sub-metrics: on T2V action prompts it tops 16/35 sub-metrics, notably character action (99.3), character expression (96.7), fore/mid/background (95.0) and viewing angle (91.1); on R2V action prompts it tops 11/38 sub-metrics, notably tone & color (100.0), camera movement (84.9), composition (83.1) and shot scale (78.2). Its action drops on these dimensions remain moderate (T2V 3.6, R2V 10.7), indicating that its Cinematic Language and character-performance strengths carry over to action content. The action performance degradation concentrates in Instruction Following and Aesthetic Quality, with two distinct L3 clusters (Table 6 in the appendix). In Instruction Following the drop is dominated by camera-work sub-metrics—camera movement (), focus (), shot scale () and composition ()—indicating that models show larger drops in camera framing and motion once the scene becomes fast and multi-agent. In Aesthetic Quality a second cluster of physical-realism sub-metrics degrades sharply: physical plausibility (), character-motion realism (), perspective (), artifacts () and clarity (). Action scenes thus reveal two concurrent degradation patterns that an aggregate score hides—models not only shift camera framing, but also show reduced physical plausibility and increased artifacts under vigorous motion. The same two clusters lead the R2V action drop, at a smaller magnitude (physical plausibility , artifacts ); we report this contrast descriptively, as the two tasks differ in prompt and sample distribution.

Top models already generalize across visual style. On R2V, the leading group scores almost identically on live-action and animation (Figure 12, right; Seedance 2.0 vs. , gap ), indicating that rendering style is no longer a constraint at the frontier. Style generalization instead shows wider variation in the mid-tier: several mid-ranked models score markedly lower on animation than live-action (Kling 3.0 Omni , ), and the loss again concentrates in Instruction Following (Kling 3.0 Omni ) and Aesthetic Quality, echoing the action analysis: the two harder distributions, action and animation, both challenge models on the same two axes.

Figure 13: T2V (left, 9 models) and R2V (right, 7 models): dialogue (hatched) vs. action (solid) mean scores (0–100) per model, ordered by dialogue score; red numbers give the drop (action dialogue). T2V: action , dialogue . R2V: action , dialogue .

5.10 R2V: Reference Types

R2V prompts condition on one of three reference types: scene (232 prompts), prop (209) and character (213). Figure 14 reports the reference-fidelity (visual-following) ranking for each type. The visual-following leader differs from the overall leader: HappyHorse 1.0 tops scene- and character-fidelity, while HappyHorse 1.1 leads on prop-fidelity, indicating that raw reference fidelity and overall film-quality are distinct capabilities. Across the three types, scene-space and character-appearance fidelity show a similar ranking order, whereas prop-reference fidelity reveals a notably different model ordering and a much tighter spread, confirming that current models still struggle to faithfully preserve a referenced prop even when scene and character references are respected.

Figure 14: R2V visual-following (reference-fidelity) ranking by reference type (columns: scene / prop / character).

5.11 Cinematic Language Related L3 Landscape

Finally, we zoom in on the L3 sub-metrics most directly tied to film production craft. Figure 15 collects 10 sub-metrics from three L2 groups across the two tasks: Cinematic Language (camera movement, shot scale, focus, viewing angle, tone & color, composition), editing (camera-work appeal, editing fluency), and performance (action performance, emotional performance). Across both tasks, models score highest on tone & color (T2V 95.0, R2V 91.2) and lowest on action performance (50.8 / 46.8) and camera-work appeal (54.8 / 57.5). The performance group is uniformly the lowest-scoring L2 cluster in both tasks, while Cinematic Language dimensions show the widest T2V–R2V gap (e.g., camera movement 69.5 vs. 55.6), indicating that reference conditioning amplifies the spread on camera-work execution.

Figure 15: Cinematic Language L3 sub-metric cross-model means (0–100), T2V (hatched) vs. R2V (solid), sorted by descending score. Bars are colored by L2 group: Cinematic Language (blue), editing (red), performance (purple).

5.12 Summary of Findings

Taken together, the analyses above distill into five findings that hold across the field. (i) No saturation; models differ most on Cinematic Language sub-metrics: no model nears the ceiling on either task (tops of 88.93 / 86.66), and camera movement, focus, shot scale and viewing angle show the largest inter-model spread, far more so under R2V (camera-movement variance vs. ), while temporal and audio dimensions show near-uniform scores. (ii) Dynamic aesthetics are the lowest-scoring sub-metrics across the field: action performance, camera-work a

正文较长,站内已展示前半部分。阅读完整原文
Hugging Face多模态开源/仓库视频论文/研究评测/基准
阅读原文arxiv.org