This is great news showing the power that agentic harness brings in the world of finance.
Primer scored 79.1% on BigFinanceBench, a 928-question finance benchmark, using GPT-5.5, the same model that alone scores 55.8%.
BigFinanceBench measures whether AI can complete real financial-analyst tasks, retrieval, calculations, modeling, and reasoning, not just answer finance trivia. You get asked to split a private equity fund's profit between its investors and its managers, or rebuild divisional numbers after a reporting change.
GPT-5.5 leads all models at 55.8%, and inside Primer's harness the same model hits 79.1%.
Primer is a finance-focused AI agent system that wraps a general model like GPT-5.5 with specialist retrieval, calculation tools, and workflow logic to perform analyst tasks more reliably.
Most of the gain, 16.3 points, came from retrieval, because the agent goes and finds the right line in the actual filing instead of reasoning over whatever the benchmark hands it.
The rest, 9.3 points, came from calculation, because Primer already knows how a company defines its own metrics, so it excludes a credit line announced six weeks after the quarter closed rather than counting it as liquidity.