GPT-6 Astra delivered Perplexity’s strongest WANDR result yet
13.5% above Fable 5.1 while costing 6.1% less per task.
This benchmark WANDR is quite unusual because it tests wide-and-deep research capability of models: finding large sets of qualifying entities, then backing every requested fact with checkable evidence.
WANDR itself contains 500 public tasks requiring 170,495 source-backed records, so incomplete research is directly penalized even when the facts an agent did find are correct.
Its scoring tracks both precision and completion, with stricter hard scores requiring an entire requested branch to be correct before receiving full credit.
So the 0.682 result points to a substantial gain on long, evidence-heavy research work where an agent must keep finding, checking, and organizing information at scale.