Ethan Mollick@emollick
52AI 编辑部评分,满分 100
2026-07-30 23:15· 25分钟前
跳到正文
AI 摘要

随着测试前沿AI的基准变得越来越复杂,我们正在失去基准测试最重要的方面之一:与人类的比较。 经过验证的基准需要有人类(最好是多个人类)基线。这做起来越来越难且成本高昂,但很重要。

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans

Validated benchmarks need to have human (ideally multiple humans) baselines. It is increasingly hard & pricey to do, but important

Ethan Mollick · @emollick · X·2026-07-30 23:15·25分钟前
在 X 看原推· x.com
AI 摘要

随着测试前沿AI的基准变得越来越复杂,我们正在失去基准测试最重要的方面之一:与人类的比较。 经过验证的基准需要有人类(最好是多个人类)基线。这做起来越来越难且成本高昂,但很重要。

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans

Validated benchmarks need to have human (ideally multiple humans) baselines. It is increasingly hard & pricey to do, but important

在 X 查看原推x.com