NARU:面向日本超长视频的叙事演变与文化细微理解基准

HuggingFace Daily Papers(社区热门论文)·2026-08-13 08:00·14天前
AI 导读

NARU 是一个评估日本长视频中叙事演变与文化推理能力的基准,包含 155 个视频、总计 146.8 小时,以及 1,481 个问题,覆盖四个叙事维度和五个文化维度。其构建采用基于分层记忆的标注流程,并经过 68 名母语者的两轮验证。对八种模型配置的评估显示,模型在长程叙事整合与文化推理方面仍存在明显局限。

HuggingFace Daily Papers(社区热门论文)
52AI 编辑部评分,满分 100

NARU:面向日本超长视频的叙事演变与文化细微理解基准

2026-08-13 08:00· 14天前
AI 导读

NARU 是一个评估日本长视频中叙事演变与文化推理能力的基准,包含 155 个视频、总计 146.8 小时,以及 1,481 个问题,覆盖四个叙事维度和五个文化维度。其构建采用基于分层记忆的标注流程,并经过 68 名母语者的两轮验证。对八种模型配置的评估显示,模型在长程叙事整合与文化推理方面仍存在明显局限。

Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org