# 当人工智能基准测试陷入停滞时：一项关于基准测试饱和度的系统性研究

- 来源：Hacker News 热门（buzzing.cc 中文翻译）
- 作者：doppp
- 发布时间：2026-08-05 19:19
- AIHOT 分数：39
- AIHOT 链接：https://aihot.virxact.com/items/cmsg0eo3m00uvrolgfkz3bbre
- 原文链接：https://arxiv.org/abs/2602.16763

## AI 摘要

一项研究分析了60个语言模型基准测试的饱和度，发现近半数基准已出现饱和，且饱和率随基准年龄增长而上升。研究还发现，基准对饱和的抵抗力受专家策划影响，而非公开测试数据。结果表明，设计选择可延长基准寿命，并为更持久的评估方法提供参考。

## 正文

Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks quickly "saturate", making it difficult to differentiate models and diminishing their long-term value. In this study, we define benchmark saturation and analyze it across 60 language model benchmarks using 14 properties that relate to saturation. We find that nearly half of the our benchmarks exhibit saturation, with rates increasing with age. Further, we find that resilience to saturation is impacted by expert-curation, not by public test data. Our results suggest that design choices can extend benchmark longevity and inform more durable evaluation approaches.
