论文:LLM 排行榜名次很大程度由评测配置决定,gemma4-31b 得分可在 31% 到 89% 间波动

Rohan Paul · @rohanpaul_ai · X·2026-09-01 05:58·10小时前
AI 导读

一篇论文在固定 12 个模型和 3,679 个问题的前提下,只改变提示词格式、选项顺序、打分方式等常规评测设置,结果排名大幅波动。gemma4-31b 得分在 31% 到 89% 之间,12 个模型中有 4 个至少在一种有效设置下排到第一;相邻模型平均 95.7% 的差距来自评测设置改变时会翻转答案的题目,最大不稳定来源是打分方式,即生成答案还是选最高似然选项。

Rohan Paul@rohanpaul_ai
60AI 编辑部评分,满分 100

论文:LLM 排行榜名次很大程度由评测配置决定,gemma4-31b 得分可在 31% 到 89% 间波动

2026-09-01 05:58· 10小时前
AI 导读

一篇论文在固定 12 个模型和 3,679 个问题的前提下,只改变提示词格式、选项顺序、打分方式等常规评测设置,结果排名大幅波动。gemma4-31b 得分在 31% 到 89% 之间,12 个模型中有 4 个至少在一种有效设置下排到第一;相邻模型平均 95.7% 的差距来自评测设置改变时会翻转答案的题目,最大不稳定来源是打分方式,即生成答案还是选最高似然选项。

LLM rankings can be created by evaluation choices as much as model differences, so one setup should never decide the leaderboard.

The paper keeps the models and 3,679 questions fixed, then changes only ordinary evaluation choices such as prompt format, option order, and scoring method.

Those choices move the results a lot.

gemma4-31b scores anywhere from 31% to 89%, and 4 of the 12 models reach rank 1 under at least 1 valid setup.

Even more telling, 95.7% of the average gap between neighboring models comes from questions whose answers change when the evaluation setup changes.

The biggest source of instability is how answers are scored: generating an answer versus choosing the highest-likelihood option.

来源:Rohan Paul· x.com