Ethan Mollick · @emollick · X·2026-09-08 12:56·17分钟前
AI 导读

Ethan Mollick 批评 GDPval-AA 是糟糕的基准:AI 回答主观判断题,且由其他 AI 而非人类专家评分,缺乏有效性却仍被 AI 实验室引用。引用推文指出 Artificial Analysis 坚持 GPT-6 在该基准上真实任务得分低于 Kimi K3,而 OpenAI 对此保持沉默,或与测试评分方式有关。

Ethan Mollick@emollick
39AI 编辑部评分,满分 100
2026-09-08 12:56· 17分钟前
AI 导读

Ethan Mollick 批评 GDPval-AA 是糟糕的基准:AI 回答主观判断题,且由其他 AI 而非人类专家评分,缺乏有效性却仍被 AI 实验室引用。引用推文指出 Artificial Analysis 坚持 GPT-6 在该基准上真实任务得分低于 Kimi K3,而 OpenAI 对此保持沉默,或与测试评分方式有关。

Broken record here, but GDPval-AA is a very bad benchmark. 1) AIs give answers to the public questions of GDPval, which involve subjective judgement 2) Other AIs judge the answer (in the real GDPval, it was human experts) It has no validity & yet even the AI labs keep citing it.

ChrisSo, one thing that still puzzles me is that Artificial Intelligence Analysis has confidently stood by their score that GPT-6 scores worse on real world tasks th...

来源:Ethan Mollick· x.com