# Artificial Analysis 评测 MiniCPM5-2B：4B 以下开源权重模型智能指数第一

- 来源：Artificial Analysis (@ArtificialAnlys)
- 发布时间：2026-09-07 21:36
- AIHOT 分数：55
- AIHOT 链接：https://aihot.virxact.com/items/cmtrb75yp01j2roa42e1t9sg5
- 原文链接：https://x.com/ArtificialAnlys/status/2096955784592797998

## AI 摘要

Artificial Analysis 公布 MiniCPM5-2B 在 Intelligence Index v4.2 得 15 分，为 4B 总参数以下开源权重模型最高分，仅次于约 3 倍参数的 Ling 3.0 Tiny（16）。

## 正文

OpenBMB's MiniCPM5-2B scores 15 on the Artificial Analysis Intelligence Index v4.2, the highest of any open weights model under 4B total parameters

OpenBMB (@OpenBMB) is the open-source AI group behind the MiniCPM series of efficient small models. MiniCPM5-2B is a 2.6B parameter dense reasoning model with text input and output, released under Apache 2.0.

Scoring 15 on the Intelligence Index, MiniCPM5-2B sits one point behind Ling 3.0 Tiny (16), which has ~3x the total parameters. Among open weights models under 4B total parameters, the next best score is Granite 4.2 3B (11).

Key results:

➤ The highest Intelligence Index of any open weights model under 4B total parameters, setting a new Pareto-optimal point on Intelligence vs. Total Parameters: Its score of 15 is 4 points clear of Granite 4.2 3B (11). With 2.6B total parameters, it is 1 point ahead of Qwen3.5 4B (Reasoning, 14, estimated) with 44% fewer parameters, and level with Qwen3.5 9B (Reasoning, 15, estimated) at roughly 4x its size. As a dense model, its size advantage is in memory footprint rather than active-parameter compute.

➤ Strong agentic performance at this size: Its GDPval-AA v2 Elo of 831 leads <4B models, and on τ³-Banking it is joint-first with Ling 3.0 Tiny at 21%, compared to 8% for the next best model, Granite 4.2 8B. On AA-Briefcase, it placed second among the measured models in the comparison set with an Elo of 438, above Granite 4.2 8B (324) and just below Ling 3.0 Tiny (485).

➤ Knowledge, coding and long context are where it gives ground: MiniCPM5-2B places 7th in the set on Humanity's Last Exam (9%, behind Gemma 4 12B (Reasoning) at 16%), 8th on Terminal-Bench v2.1 (9%, behind Qwen3.5 9B (Reasoning) at 29%) and scores 0% on CritPt. On SciCode it is second of the five measured models at 26%, behind Granite 4.2 8B (31%). On AA-LCR v1.1 it scores 59%, 5th in the set, one point behind Ling 3.0 Tiny (60%). On GDP.pdf, our new professional document reasoning evaluation, it passes 1% of tasks outright, behind gpt-oss-20b (high) at 2%.

➤ Its AA-Omniscience score of -12 is earned by abstaining from answering rather than accuracy: MiniCPM5-2B attempts only 29% of AA-Omniscience questions, giving it a Non-Hallucination Rate of 78%. Its accuracy of 8% is a point below Ling 3.0 Tiny (9%) and half that of Qwen3.5 9B (Reasoning, 16%). Peers that attempt far more questions are penalized heavily, with Qwen3.5 9B (Reasoning) at -53 and gpt-oss-20b (high) at -63.

➤ It is token-efficient for a reasoning model: MiniCPM5-2B used 19k output tokens per Intelligence Index task, joint-lowest in the comparison model set with Granite 4.2 3B (19k). Ling 3.0 Tiny spends 56k, roughly 3x as many, for 1 more index point.

Additional model details:

➤ Parameters: 2.6B (dense)

➤ Context window: 131k tokens

➤ Input modalities: Text only

➤ License: Apache 2.0
