量化 ASR 模型基准优化:高分开源模型在音频不充分时仍复现基准参考文本

HuggingFace Daily Papers(社区热门论文)·2026-08-20 08:00·5天前
AI 导读

一项新研究提出量化自动语音识别(ASR)模型基准优化程度的方法,聚焦音频不足以确定参考转写文本的情形。研究发现,得分最高的开源模型即使在相关音频矛盾、被遮蔽或含混时,仍会逐字输出基准参考文本片段,且该行为可通过低秩线性引导或简单在片段末尾追加音频进行因果操控。结果表明,高性能模型存在基准条件化行为,可能虚增基准分数而不反映通用转写能力的提升。

HuggingFace Daily Papers(社区热门论文)
50AI 编辑部评分,满分 100

量化 ASR 模型基准优化:高分开源模型在音频不充分时仍复现基准参考文本

2026-08-20 08:00· 5天前
AI 导读

一项新研究提出量化自动语音识别(ASR)模型基准优化程度的方法,聚焦音频不足以确定参考转写文本的情形。研究发现,得分最高的开源模型即使在相关音频矛盾、被遮蔽或含混时,仍会逐字输出基准参考文本片段,且该行为可通过低秩线性引导或简单在片段末尾追加音频进行因果操控。结果表明,高性能模型存在基准条件化行为,可能虚增基准分数而不反映通用转写能力的提升。

Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org