# LVSum 基准：评估多模态大模型的长视频时间感知摘要能力

- 来源：Apple Machine Learning Research（RSS）
- 发布时间：2026-07-20 08:00
- AIHOT 分数：49
- AIHOT 链接：https://aihot.virxact.com/items/cmrtiqww63mwvbitlt2z8my54
- 原文链接：https://machinelearning.apple.com/research/lvsum-video-summarization

## AI 摘要

Apple 推出 LVSum，一个含细粒度时间对齐的人工标注长视频摘要基准，包含 13 个领域的 72 个视频（平均时长 16 分钟），每个视频配有最多 10 条含时间引用的人类摘要。评估显示，转录文本对摘要质量的贡献远大于视觉帧，当前多模态大模型在时间定位、指令遵循和跨模态一致性上存在系统性缺陷，与人类摘要仍有显著差距。

## 正文

Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded. We introduce LVSum, a human-annotated benchmark for evaluating long-form video summarization with fine-grained temporal alignment. LVSum comprises 72 diverse videos spanning 13 domains with an average duration of 16 minutes, each annotated with up to 10 human-generated summaries containing temporal references. We conduct a comprehensive evaluation of leading proprietary and open-source MLLMs using newly introduced LLM-based metrics for content relevance and modality coherence, alongside standard automatic metrics. Our experiments reveal three key findings: (1) transcripts contribute substantially more to summarization quality than visual frames alone, (2) a significant performance gap persists between model-generated and human-written summaries, and (3) current MLLMs exhibit systematic weaknesses in temporal grounding, instruction adherence, and cross-modal coherence.

Related readings and updates.

Incentivizing Temporal-Awareness in Egocentric Video Understanding Models

Multimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings where reasoning depends on the correct ordering and evolution of events. This deficiency stems in part from training objectives that fail to explicitly reward temporal reasoning and instead rely on frame-level spatial shortcuts. To address this limitation, we propose…

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding

Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand temporally unfolding narratives in videos remains underexplored. True narrative understanding requires grounding who is doing what, when, and where, maintaining coherent entity representations across dynamic visual and temporal contexts. We introduce NarrativeTrack, the first benchmark to evaluate narrative…

Discover opportunities in Machine Learning.
