Perplexity 实测 GPT-6 Astra 取得 WANDR 最高分 0.682

Rohan Paul · @rohanpaul_ai · X·2026-09-04 11:09·5分钟前
AI 导读

Perplexity 在 WANDR 基准上评测 GPT-6 Astra,得分 0.682、每任务成本 $11.98,为其测试过的模型中最高分;比 Fable 5.1 高 13.5% 且成本低 6.1%,比 Opus 5 高 27.0% 且成本高 3.3%。

Rohan Paul@rohanpaul_ai
50AI 编辑部评分,满分 100

Perplexity 实测 GPT-6 Astra 取得 WANDR 最高分 0.682

2026-09-04 11:09· 5分钟前
AI 导读

Perplexity 在 WANDR 基准上评测 GPT-6 Astra,得分 0.682、每任务成本 $11.98,为其测试过的模型中最高分;比 Fable 5.1 高 13.5% 且成本低 6.1%,比 Opus 5 高 27.0% 且成本高 3.3%。

GPT-6 Astra delivered Perplexity’s strongest WANDR result yet

13.5% above Fable 5.1 while costing 6.1% less per task.

This benchmark WANDR is quite unusual because it tests wide-and-deep research capability of models: finding large sets of qualifying entities, then backing every requested fact with checkable evidence.

WANDR itself contains 500 public tasks requiring 170,495 source-backed records, so incomplete research is directly penalized even when the facts an agent did find are correct.

Its scoring tracks both precision and completion, with stricter hard scores requiring an entire requested branch to be correct before receiving full credit.

So the 0.682 result points to a substantial gain on long, evidence-heavy research work where an agent must keep finding, checking, and organizing information at scale.

PerplexityWe evaluated GPT-6 Astra on WANDR. It scored 0.682 at $11.98 per task, the highest score of any model we tested. GPT-6-Astra scored 13.5% higher than Fable 5.1 ...

来源:Rohan Paul· x.com