HuggingFace Daily Papers(社区热门论文)
57AI 编辑部评分,满分 100

AVA-Encoder:面向智能体的视频表示学习框架

2026-08-12 08:00· 1天前
AI 导读

AVA-Encoder 提出智能体视频自编码框架,将视频转化为知识图谱表示再重建回视频,以文本梯度优化驱动策略训练。该方法在评测中较最强外部基线提升 20.7 个百分点,伪训练策略在系统提示词 token 减少 74.3% 的情况下仍优于人工调优策略。研究团队已开源完整框架、重建基准及首个高质量电影知识图谱数据集。

Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations via agentic auto-encoding. AVA-Encoder transforms a video into a knowledge graph (KG) representation and then reconstructs it back into video. Its hierarchy and state nodes store structured text, while a linked asset layer holds generated images, audio, and video. Typed edges preserve the relations between these text descriptions and assets in a form that agents can easily understand, query, and edit. The video reconstruction differences drive a textual-gradient optimization framework, which expresses evaluation feedback as natural-language update directions for Data-Independent Encoding Policy Pseudo-Training in the outer loop and optional Data-Dependent KG Representation Refinement in the test-time inner loop. Extensive experiments show that AVA-Encoder improves by 20.7 percentage points over the strongest external baseline. In the controlled policy-only setting, its pseudo-trained shot-level Agentic Video Encoder policy also outperforms a carefully human-tuned policy while using 74.3% fewer system-prompt tokens. We release the complete AVA-Encoder framework, a reliable agentic video reconstruction benchmark, and the first dataset of high-quality film KG representations.

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org

AVA-Encoder:面向智能体的视频表示学习框架

HuggingFace Daily Papers(社区热门论文)·2026-08-12 08:00·1天前
AI 导读

AVA-Encoder 提出智能体视频自编码框架,将视频转化为知识图谱表示再重建回视频,以文本梯度优化驱动策略训练。该方法在评测中较最强外部基线提升 20.7 个百分点,伪训练策略在系统提示词 token 减少 74.3% 的情况下仍优于人工调优策略。研究团队已开源完整框架、重建基准及首个高质量电影知识图谱数据集。

原文 · 保持原样,未翻译

Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations via agentic auto-encoding. AVA-Encoder transforms a video into a knowledge graph (KG) representation and then reconstructs it back into video. Its hierarchy and state nodes store structured text, while a linked asset layer holds generated images, audio, and video. Typed edges preserve the relations between these text descriptions and assets in a form that agents can easily understand, query, and edit. The video reconstruction differences drive a textual-gradient optimization framework, which expresses evaluation feedback as natural-language update directions for Data-Independent Encoding Policy Pseudo-Training in the outer loop and optional Data-Dependent KG Representation Refinement in the test-time inner loop. Extensive experiments show that AVA-Encoder improves by 20.7 percentage points over the strongest external baseline. In the controlled policy-only setting, its pseudo-trained shot-level Agentic Video Encoder policy also outperforms a carefully human-tuned policy while using 74.3% fewer system-prompt tokens. We release the complete AVA-Encoder framework, a reliable agentic video reconstruction benchmark, and the first dataset of high-quality film KG representations.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org