# 前沿AI基准测试正失去人类基线对比

- 来源：Ethan Mollick (@emollick)
- 发布时间：2026-07-30 23:15
- AIHOT 分数：52
- AIHOT 链接：https://aihot.virxact.com/items/cms7o4yj60nxiro2ef49c0aea
- 原文链接：https://x.com/emollick/status/2082847594519130480

## AI 摘要

随着测试前沿AI的基准变得越来越复杂，我们正在失去基准测试最重要的方面之一：与人类的比较。

经过验证的基准需要有人类（最好是多个人类）基线。这做起来越来越难且成本高昂，但很重要。

## 正文

As the benchmarks that test frontier AI on get more complex， we are losing one of the most important aspects of benchmarking： comparisons to humans

Validated benchmarks need to have human （ideally multiple humans） baselines. It is increasingly hard &amp； pricey to do， but important
