# Agnes 2.5 Pro Beta 智能指数跃升至 49

- 来源：Artificial Analysis (@ArtificialAnlys)
- 发布时间：2026-08-28 10:04
- AIHOT 分数：52
- AIHOT 链接：https://aihot.virxact.com/items/cmtcbvt8o0208ro5gxph1qvsz
- 原文链接：https://x.com/ArtificialAnlys/status/2093157744471851476

## AI 摘要

Agnes AI 的 Agnes 2.5 Pro Beta 在 Artificial Analysis 智能指数上得分 49，较 Alpha 版提升 9 分，主要得益于智能体能力的大幅增长（Agentic Index 从 25 升至 44），但输出 token 消耗约为前代的两倍。

## 正文

Agnes AI's Agnes 2.5 Pro Beta scores 49 on the Artificial Analysis Intelligence Index, up 9 points from Agnes 2.5 Pro Alpha, driven by large agentic gains but using ~2x the output tokens

Agnes AI (@agnesai_sapiens) is a Singapore-based AI lab that trains full-modality foundation models and offers them through a free omni-modal API.

At 49 on the Intelligence Index, Agnes 2.5 Pro Beta moves Agnes from mid-pack to the frontier-adjacent tier, just behind Gemini 3.5 Flash (high, 52) and GPT-5.6 Luna (max, 52) and ahead of MiniMax-M3 (45).

Key results:

➤ Agnes 2.5 Pro Beta scores 49 on the Artificial Analysis Intelligence Index, a 9-point jump from Agnes 2.5 Pro Alpha (40). This places it just behind Gemini 3.5 Flash (high, 52) and GPT-5.6 Luna (max, 52), and ahead of MiniMax-M3 (45).

➤ Agentic capabilities drive the jump. The Artificial Analysis Agentic Index rises from 25 to 44, just behind Gemini 3.7 Flash (high, 45) and ahead of Gemini 3.5 Flash (high, 40) and MiniMax-M3 (36). τ³-Banking nearly triples from 12% to 36%, and GDPval-AA v2 rises from an Elo of 1171 to 1456 against a human baseline of 1000.

➤ Frontier reasoning evaluations improves more modestly. Humanity's Last Exam rises from 34% to 38%, GPQA Diamond from 88% to 91%, and CritPt from 11% to 16%.

➤ The AA-Omniscience improvement from -25 to -11 comes from abstention, not increased accuracy. Agnes 2.5 Pro Beta attempts only 45% of questions against Agnes 2.5 Pro Alpha's 94%, cutting the hallucination rate from 88% to 33%, but AA-Omniscience Accuracy halves from 33% to 17%.

➤ The intelligence gain required ~2x as many tokens from its predecessor. Agnes 2.5 Pro Beta uses 50k output tokens per Intelligence Index task, more than double Agnes 2.5 Pro Alpha's 24k.

Additional model details:

➤ Context window: 1M tokens

➤ Max output tokens: 65k

➤ Input modalities: Text and image

➤ Pricing: $0.10 / $0.30 / $0.01 per 1M input / output / cache hit tokens

➤ Availability: Agnes AI first-party API
