# Thinking Machines 发布生产级语言模型 Inkling，成美国实验室最强开源模型

- 来源：Artificial Analysis (@ArtificialAnlys)
- 发布时间：2026-07-16 02:32
- AIHOT 分数：71
- AIHOT 链接：https://aihot.virxact.com/items/cmrmfi33l00rvbimr1h0zczg0
- 原文链接：https://x.com/ArtificialAnlys/status/2077461348079001716

## AI 摘要

Thinking Machines 发布首个生产级语言模型 Inkling，总参数量 975B，激活 41B，支持文本、图像和音频多模态输入。该模型在 Artificial Analysis Intelligence Index 上以 41 分成为美国实验室最强开源模型，超越 Nemotron 3 Ultra（38 分）和 Gemma 4 31B（29 分）。

## 正文

Thinking Machines has released Inkling， the new leading U.S. open weights model， debuting at 41 on the Artificial Analysis Intelligence Index

@thinkymachines has previously released research previews of models and this is their first production language model release. The model is 975B total parameters， has 41B active parameters， and accepts text， image， and audio input modalities. The model is accessible via Thinking Machines' Tinker platform API （256K context window） and weights are available on HuggingFace （1M context window）.

Key results：

➤ Inkling debuts at 41 on the Artificial Analysis Intelligence Index， making it the leading open weights release from a U.S. lab. Inkling scores 3 points higher on the Intelligence Index （41） than the previous leading U.S. open weights model， Nemotron 3 Ultra （38）， and also beats Gemma 4 31B （29） and gpt-oss-120b （24）

➤ Inkling stands out on agentic performance. It scores higher than both Kimi K.6 and DeepSeek v4 Flash on both GDPval-AA v2 and τ3-Banking： Inkling scores an Elo of 1238 on GDPval-AA v2， higher than Kimi K.6 （1190） and DeepSeek v4 Flash max （1189） and scores 24% on τ3-Banking， higher than Kimi K2.6 （21%） and just above DeepSeek v4 Flash max （23%）

➤ Inkling is token efficient compared to open weights leaders. Inkling averages 25K output tokens per Intelligence Index task compared to 43K， 38K and 37K by GLM-5.2 （max）， Kimi K2.6 and DeepSeek v4 Pro （max） respectively

➤ Inkling natively supports multimodal input， a key differentiator among open weights models. Inkling accepts text， image， and audio input modalities. Images and videos are encoded via a hierarchical patch encoder and audio via discrete token encoding， with all modalities projected into a shared hidden space and processed jointly by the decoder

Additional model details：

➤ Type： Open weights reasoning model

➤ Size： 975B （41B active） parameters

➤ Input modalities： Text， image， and audio （text output）

➤ Context window： 256K tokens on Tinker， open weights model supports 1M

➤ Pricing per 1M tokens （64K context window）： $1.87 input / $0.374 cached / $4.68 output

➤ Pricing per 1M tokens （256K context window）： $3.74 input / $0.748 cached / $9.36 output
