OpenBMB's MiniCPM5-2B scores 15 on the Artificial Analysis Intelligence Index v4.2, the highest of any open weights model under 4B total parameters
OpenBMB (@OpenBMB) is the open-source AI group behind the MiniCPM series of efficient small models. MiniCPM5-2B is a 2.6B parameter dense reasoning model with text input and output, released under Apache 2.0.
Scoring 15 on the Intelligence Index, MiniCPM5-2B sits one point behind Ling 3.0 Tiny (16), which has ~3x the total parameters. Among open weights models under 4B total parameters, the next best score is Granite 4.2 3B (11).
Key results:
➤ The highest Intelligence Index of any open weights model under 4B total parameters, setting a new Pareto-optimal point on Intelligence vs. Total Parameters: Its score of 15 is 4 points clear of Granite 4.2 3B (11). With 2.6B total parameters, it is 1 point ahead of Qwen3.5 4B (Reasoning, 14, estimated) with 44% fewer parameters, and level with Qwen3.5 9B (Reasoning, 15, estimated) at roughly 4x its size. As a dense model, its size advantage is in memory footprint rather than active-parameter compute.
➤ Strong agentic performance at this size: Its GDPval-AA v2 Elo of 831 leads <4B models, and on τ³-Banking it is joint-first with Ling 3.0 Tiny at 21%, compared to 8% for the next best model, Granite 4.2 8B. On AA-Briefcase, it placed second among the measured models in the comparison set with an Elo of 438, above Granite 4.2 8B (324) and just below Ling 3.0 Tiny (485).
➤ Knowledge, coding and long context are where it gives ground: MiniCPM5-2B places 7th in the set on Humanity's Last Exam (9%, behind Gemma 4 12B (Reasoning) at 16%), 8th on Terminal-Bench v2.1 (9%, behind Qwen3.5 9B (Reasoning) at 29%) and scores 0% on CritPt. On SciCode it is second of the five measured models at 26%, behind Granite 4.2 8B (31%). On AA-LCR v1.1 it scores 59%, 5th in the set, one point behind Ling 3.0 Tiny (60%). On GDP.pdf, our new professional document reasoning evaluation, it passes 1% of tasks outright, behind gpt-oss-20b (high) at 2%.
➤ Its AA-Omniscience score of -12 is earned by abstaining from answering rather than accuracy: MiniCPM5-2B attempts only 29% of AA-Omniscience questions, giving it a Non-Hallucination Rate of 78%. Its accuracy of 8% is a point below Ling 3.0 Tiny (9%) and half that of Qwen3.5 9B (Reasoning, 16%). Peers that attempt far more questions are penalized heavily, with Qwen3.5 9B (Reasoning) at -53 and gpt-oss-20b (high) at -63.
➤ It is token-efficient for a reasoning model: MiniCPM5-2B used 19k output tokens per Intelligence Index task, joint-lowest in the comparison model set with Granite 4.2 3B (19k). Ling 3.0 Tiny spends 56k, roughly 3x as many, for 1 more index point.
Additional model details:
➤ Parameters: 2.6B (dense)
➤ Context window: 131k tokens
➤ Input modalities: Text only
➤ License: Apache 2.0