Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.
RA-Bench:系统评估AI生成视频检测器在真实危机事件中的表现
AI 导读
RA-Bench基准以真实视频为锚点,含17,886个视频,覆盖10类社会风险场景,由1,830个真实视频和来自4个开源、5个闭源生成器的16,056个生成片段组成。评估发现,七种传统检测器、十种零样本多模态模型及两种微调MLLM均无法在RA-Bench上稳定泛化;能误导人类的视频同样难以被现有检测器识别,且社交传播进一步降低检测可靠性。
HuggingFace Daily Papers(社区热门论文)
64
AI 编辑部评分,满分 100RA-Bench:系统评估AI生成视频检测器在真实危机事件中的表现
RA-Bench基准以真实视频为锚点,含17,886个视频,覆盖10类社会风险场景,由1,830个真实视频和来自4个开源、5个闭源生成器的16,056个生成片段组成。评估发现,七种传统检测器、十种零样本多模态模型及两种微调MLLM均无法在RA-Bench上稳定泛化;能误导人类的视频同样难以被现有检测器识别,且社交传播进一步降低检测可靠性。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org