Streaming video understanding demands direct responses from the causally observed prefix of an unfolding video. Existing systems add inference-time memory, retrieval, and compression, yet a training-free sliding-window baseline already matches them. We therefore fix a memory-free recent-window protocol and ask how far post-training alone can go. Reinforcement learning with verifiable rewards fits this regime poorly, encouraging long ``think-then-answer'' generations, while on-policy distillation (OPD) supplies dense token-level teacher supervision on student trajectories but is stable only when both models train in thinking mode. These observations lead to StreamOPD, a recipe combining verifiable streaming-video data, thinking-mode OPD, and instruct-mode deployment. It raises StreamingBench from 77.9% to 83.9%---within 0.3 points of the 9B teacher---and improves OVO-Bench excluding its hallucination-detection subtask (HLD) by 9.1 points under unchanged inference. As a teacher-privilege extension, Spatio-Temporal CueGate (ST-CueGate) aggregates cue-versus-no-cue teacher likelihood ratios into a group-relative response score that reweights OPD. It reaches 71.9% on OVO-Bench (excluding HLD) and 64.9% on Video-MME, and is the only variant that stays above the base model on all four benchmarks. Replacing the teacher with a frozen copy of the student's initial policy---on-policy self-distillation---retains most of these gains and lifts HLD to 57.0%, above both the untrained student and the 9B teacher, so abstention loss is not intrinsic to the recipe. We provide a transparent and reproducible reference for open-source streaming-video research.
StreamOPD:面向流式视频理解的后训练方案与时空线索门控
AI 导读
StreamOPD 提出免推理时记忆的后训练方案,将 StreamingBench 从 77.9% 提升至 83.9%,距 9B 教师模型仅差 0.3 个百分点。其扩展模块 ST-CueGate 在 OVO-Bench(排除幻觉检测子任务)和 Video-MME 上分别达 71.9% 和 64.9%,是唯一在所有四项基准上均高于基础模型的变体。
HuggingFace Daily Papers(社区热门论文)
52
AI 编辑部评分,满分 100StreamOPD:面向流式视频理解的后训练方案与时空线索门控
StreamOPD 提出免推理时记忆的后训练方案,将 StreamingBench 从 77.9% 提升至 83.9%,距 9B 教师模型仅差 0.3 个百分点。其扩展模块 ST-CueGate 在 OVO-Bench(排除幻觉检测子任务)和 Video-MME 上分别达 71.9% 和 64.9%,是唯一在所有四项基准上均高于基础模型的变体。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org