Thinking Machines' new Inkling Small scores 40 on the Artificial Analysis Intelligence Index, within a point of its flagship sibling Inkling with less than a third of the total and active parameters
Inkling Small is @thinkymachines second model release, arriving two weeks after Inkling launched at 41 on the Artificial Analysis Intelligence Index. Thinking Machines is the San Francisco-based AI lab founded by former @OpenAI CTO Mira Murati. Inkling Small is an open weights reasoning model with 276B total parameters (12B active MoE), text, image, and speech input, and a 256K token context window.
Key results:
➤ Inkling Small holds a similar tier of intelligence to Inkling's at less than one third of its size: 276B total parameters (12B active) vs. 975B (41B active) for Inkling. No open weights model at its size or smaller scores higher on the Intelligence Index. DeepSeek V4 Flash (max), at a similar size of 284B total (13B active) also scores 40, MiniMax-M3 reaches 44 with 23B active, while GLM-5.2 (max) reaches 51 with 40B active.
➤ Inkling Small meets or exceeds Inkling on several coding and frontier reasoning evaluations. It scores higher on Humanity's Last Exam (32% vs. 30%), GPQA Diamond (89% vs. 87%), CritPt (8% vs. 5%), and SciCode (49% vs. 46%), and achieves the same score on Terminal Bench v2.1 (55%).
➤ Inkling Small is not as strong as Inkling on Agentic tasks and factual knowledge. Inkling Small trails Inkling on τ3-Banking (15% vs. 24%), though it edges ahead on GDPval-AA v2 (1269 vs. 1237 Elo). On the AA-Omniscience Index it scores -9 vs. the flagship's positive 2, meaning incorrect answers outweigh correct ones; this is driven by lower AA-Omniscience Accuracy (31% vs. 40%) rather than hallucination: Inkling Small's Hallucination Rate is slightly lower than its larger sibling (57% vs. 63%).
➤ Inkling Small averaged ~24K output tokens per Intelligence Index task, slightly fewer than Inkling (~25K), while peers at its intelligence level averaged far more. DeepSeek V4 Flash averaged ~45K and GPT-5.4 mini (xhigh) ~78K. Running the full Intelligence Index took roughly the same number of output tokens for Inkling Small and Inkling.
Additional model details:
➤ Type: Open weights reasoning model (Apache 2.0 license)
➤ Size: 276B total parameters, 12B active (MoE)
➤ Input modalities: Text, image, and speech
➤ Output modalities: Text
➤ Context window: 256K tokens (Inkling supports 1M)
Congratulations to the team at Thinking Machines on the release!