HuggingFace Daily Papers(社区热门论文)
53AI 编辑部评分,满分 100

EffectLearner:面向真实世界视频物体移除的物体-效果推理框架

2026-08-06 08:00· 1天前
AI 导读

EffectLearner 提出一种语义推理增强的视频物体移除框架,结合基于 VLM 的 Object-Effect Reasoner 与基于 DiT 的 Video Eraser,通过结构化效果分析提示词提取效果感知上下文,并引入运动感知掩码引导与运动一致性监督。

Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org

EffectLearner:面向真实世界视频物体移除的物体-效果推理框架

HuggingFace Daily Papers(社区热门论文)·2026-08-06 08:00·1天前
AI 导读

EffectLearner 提出一种语义推理增强的视频物体移除框架,结合基于 VLM 的 Object-Effect Reasoner 与基于 DiT 的 Video Eraser,通过结构化效果分析提示词提取效果感知上下文,并引入运动感知掩码引导与运动一致性监督。

原文 · 保持原样,未翻译

Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org