IBM 新框架 BenchDrift 揭示措辞致 LLM 分数波动 74.7 个百分点

Rohan Paul · @rohanpaul_ai · X·2026-08-20 02:45·6天前
AI 导读

IBM 新论文提出 BenchDrift 审计框架,通过在不改变答案的前提下系统改写同一问题,衡量模型分数因措辞产生的波动。在 8 个模型和 3 个基准测试中,最佳与最差准确率平均差距达 74.7 个百分点,且更强模型受影响更大。即便在最高置信度区间,改写后仍有 18.5% 的正确回答丢失。

Rohan Paul@rohanpaul_ai
47AI 编辑部评分,满分 100

IBM 新框架 BenchDrift 揭示措辞致 LLM 分数波动 74.7 个百分点

2026-08-20 02:45· 6天前
AI 导读

IBM 新论文提出 BenchDrift 审计框架,通过在不改变答案的前提下系统改写同一问题,衡量模型分数因措辞产生的波动。在 8 个模型和 3 个基准测试中,最佳与最差准确率平均差距达 74.7 个百分点,且更强模型受影响更大。即便在最高置信度区间,改写后仍有 18.5% 的正确回答丢失。

A model can know the answer and still fail because you asked the same question differently.

New IBM paper introduces BenchDrift, an auditing framework for existing benchmarks, that systematically rephrases the same questions without changing their answers, then measures how much a model's score moves purely because of wording.

Across 8 models and 3 benchmarks, the gap between best-case and worst-case accuracy averaged 74.7 percentage points.

Stronger models were actually more exposed.

For every model-benchmark pair above 60% baseline accuracy, rephrasing broke more correct answers than it recovered.

And confidence did not solve this.

Even in the highest-confidence bucket, 18.5% of correct answers were lost after meaning-preserving rewording.

So if your model selection depends on small score differences, test several equivalent phrasings or report the range.

– arxiv. org/abs/2608.11694

Title: "The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance"