Rohan Paul@rohanpaul_ai
37AI 编辑部评分,满分 100
2026-08-06 20:56· 1天前
AI 导读

金融 AI 智能体系统 Primer 在 BigFinanceBench(928 道真实分析师任务题)上以 79.1% 的准确率登顶,而同一 GPT-5.5 模型裸测仅得 55.8%。16.3 分的提升来自检索(智能体从实际文件中定位数据),9.3 分来自计算(Primer 内置公司自定义指标逻辑)。该评测现已被列入 OpenAI GPT-5.6 发布页的唯一金融基准。

This is great news showing the power that agentic harness brings in the world of finance.

Primer scored 79.1% on BigFinanceBench, a 928-question finance benchmark, using GPT-5.5, the same model that alone scores 55.8%.

BigFinanceBench measures whether AI can complete real financial-analyst tasks, retrieval, calculations, modeling, and reasoning, not just answer finance trivia. You get asked to split a private equity fund's profit between its investors and its managers, or rebuild divisional numbers after a reporting change.

GPT-5.5 leads all models at 55.8%, and inside Primer's harness the same model hits 79.1%.

Primer is a finance-focused AI agent system that wraps a general model like GPT-5.5 with specialist retrieval, calculation tools, and workflow logic to perform analyst tasks more reliably.

Most of the gain, 16.3 points, came from retrieval, because the agent goes and finds the right line in the actual filing instead of reasoning over whatever the benchmark hands it.

The rest, 9.3 points, came from calculation, because Primer already knows how a company defines its own metrics, so it excludes a credit line announced six weeks after the quarter closed rather than counting it as liquidity.

Al SmallwoodPrimer tops BigFinanceBench. We ran @primerapp_ on @RogoAI's BigFinanceBench. Primer tops the leaderboard: 79.1% final-answer accuracy vs 55.8% for the best fro...

来源:Rohan Paul · x.com

Rohan Paul · @rohanpaul_ai · X·2026-08-06 20:56·1天前
AI 导读

金融 AI 智能体系统 Primer 在 BigFinanceBench(928 道真实分析师任务题)上以 79.1% 的准确率登顶,而同一 GPT-5.5 模型裸测仅得 55.8%。16.3 分的提升来自检索(智能体从实际文件中定位数据),9.3 分来自计算(Primer 内置公司自定义指标逻辑)。该评测现已被列入 OpenAI GPT-5.6 发布页的唯一金融基准。

This is great news showing the power that agentic harness brings in the world of finance.

Primer scored 79.1% on BigFinanceBench, a 928-question finance benchmark, using GPT-5.5, the same model that alone scores 55.8%.

BigFinanceBench measures whether AI can complete real financial-analyst tasks, retrieval, calculations, modeling, and reasoning, not just answer finance trivia. You get asked to split a private equity fund's profit between its investors and its managers, or rebuild divisional numbers after a reporting change.

GPT-5.5 leads all models at 55.8%, and inside Primer's harness the same model hits 79.1%.

Primer is a finance-focused AI agent system that wraps a general model like GPT-5.5 with specialist retrieval, calculation tools, and workflow logic to perform analyst tasks more reliably.

Most of the gain, 16.3 points, came from retrieval, because the agent goes and finds the right line in the actual filing instead of reasoning over whatever the benchmark hands it.

The rest, 9.3 points, came from calculation, because Primer already knows how a company defines its own metrics, so it excludes a credit line announced six weeks after the quarter closed rather than counting it as liquidity.

Al SmallwoodPrimer tops BigFinanceBench. We ran @primerapp_ on @RogoAI's BigFinanceBench. Primer tops the leaderboard: 79.1% final-answer accuracy vs 55.8% for the best fro...

来源:Rohan Paul· x.com