TRACES:首个衡量探索式AI的基准

Rohan Paul · @rohanpaul_ai · X·2026-08-20 04:25·6天前
AI 导读

Apodex 推出 TRACES,号称全球首个衡量“探索式AI”的基准,将AI评测从“已知答案的检索”转向“无答案问题的完整调查过程”。该基准由创始人陈天桥提出,定义了六项能力,并发布“探索式智能”定义、评估标准及求解者招募。

Rohan Paul@rohanpaul_ai
41AI 编辑部评分,满分 100

TRACES:首个衡量探索式AI的基准

2026-08-20 04:25· 6天前
AI 导读

Apodex 推出 TRACES,号称全球首个衡量“探索式AI”的基准,将AI评测从“已知答案的检索”转向“无答案问题的完整调查过程”。该基准由创始人陈天桥提出,定义了六项能力,并发布“探索式智能”定义、评估标准及求解者招募。

This is a good example of why final-answer accuracy can hide bad agent behaviour.

Most AI benchmarks measure whether a model can reach a known answer.

Apodex introduced TRACES 🧭, the world's first benchmark for measuring discoverative AI.

A shift from benchmarking models on solved problems to evaluating systems that can investigate consequential problems under evidence, tools, and verification.

TRACES says AI discovery should be evaluated as an entire investigation, not as a single final answer.

This can separate a high-scoring outcome from the quality of the process that produced it, including whether errors were repaired and claims were grounded.

ApodexMost AI benchmarks test retrieval — can a model find the known answer? However, the hardest problems in science require discovery, can a system earn an answer n...

来源:Rohan Paul· x.com