elvis@omarsar0
42AI 编辑部评分,满分 100

IBM 新研究:基准分数部分源于措辞而非模型

2026-08-16 01:12· 3小时前
AI 导读

IBM 新研究提出 BenchDrift,通过生成保持答案不变但沿语言、指代、语用和结构轴变化的基准问题变体,衡量模型正确率翻转的频率。研究发现措辞敏感性不随模型能力提升而消失,反而改变方向:弱模型从改写中获益多于损失,强模型损失远超获益,因此基准排名靠前的模型其分数最依赖措辞。八款模型在 GSM8K、MMLU 和 MATH-Hard 上对哪些改写代价最高基本达成一致。

Interesting new research from IBM.

If you pick models from benchmark deltas, some of that delta belongs to the phrasing rather than the model.

BenchDrift generates meaning-preserving variations of benchmark problems along linguistic, referential, pragmatic, and structural axes, holding the answer fixed, then measures how often correctness flips.

Phrasing sensitivity does not fade as models improve. It changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain, so the top models on a benchmark are the ones whose scores depend most on the wording they happened to receive.

Fragility also belongs to the rephrasing. Across eight models on GSM8K, MMLU, and MATH-Hard, they largely agree on which rephrasings cost the most correct answers even while differing in how much they drift overall.

Rephrasing breaks answers models were confident about, whether the problem gets shorter or longer.

Paper: https://arxiv.org/abs/2608.11694

Track more trending AI papers in our academy: https://academy.dair.ai/

来源:elvis · x.com

IBM 新研究:基准分数部分源于措辞而非模型

elvis · @omarsar0 · X·2026-08-16 01:12·3小时前
AI 导读

IBM 新研究提出 BenchDrift,通过生成保持答案不变但沿语言、指代、语用和结构轴变化的基准问题变体,衡量模型正确率翻转的频率。研究发现措辞敏感性不随模型能力提升而消失,反而改变方向:弱模型从改写中获益多于损失,强模型损失远超获益,因此基准排名靠前的模型其分数最依赖措辞。八款模型在 GSM8K、MMLU 和 MATH-Hard 上对哪些改写代价最高基本达成一致。

Interesting new research from IBM.

If you pick models from benchmark deltas, some of that delta belongs to the phrasing rather than the model.

BenchDrift generates meaning-preserving variations of benchmark problems along linguistic, referential, pragmatic, and structural axes, holding the answer fixed, then measures how often correctness flips.

Phrasing sensitivity does not fade as models improve. It changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain, so the top models on a benchmark are the ones whose scores depend most on the wording they happened to receive.

Fragility also belongs to the rephrasing. Across eight models on GSM8K, MMLU, and MATH-Hard, they largely agree on which rephrasings cost the most correct answers even while differing in how much they drift overall.

Rephrasing breaks answers models were confident about, whether the problem gets shorter or longer.

Paper: https://arxiv.org/abs/2608.11694

Track more trending AI papers in our academy: https://academy.dair.ai/

来源:elvis· x.com