New Google Paper says financial deep-research agents are far better at reconstructing the past than anticipating what comes next.
Across 17 baselines plus FinanceHarness, every model they tested stayed below 40% overall on 400 expert-annotated questions.
The harness matters: with the same Qwen3.6-27B backbone, moving from a simple search loop to the full finance-oriented tool and workflow stack raised the overall score from 25.3% to 32.4%.
Extra training barely changed that result, adding only 0.4 percentage points after Group Relative Policy Optimization.
So the main bottleneck is no longer just search, citation, or report structure.
Financial research agents need better causal and scenario reasoning, because a cleaner evidence pipeline does not automatically produce better forward-looking judgment.
- arxiv. org/abs/2607.27853
Title: "FinanceHarness: Autonomous Financial Deep Research Framework"