# OpenAI 将在 Cerebras 上推出 GPT-5.6 Sol，推理速度达 750 tokens/s

- 来源：Chubby♨️ (@kimmonismus)
- 发布时间：2026-07-06 15:39
- AIHOT 分数：73
- AIHOT 链接：https://aihot.virxact.com/items/cmr8xovv201gbsllstg6owgxe
- 原文链接：https://x.com/kimmonismus/status/2074035567906426886

## AI 摘要

OpenAI 将于 7 月在 Cerebras 上推出 GPT-5.6 Sol，推理速度达 750 tokens/s。据 Bleys Goodson 估算，模型横跨 70–100 片 Cerebras 晶圆，每片约承载一层，总参数量约 3T，活跃参数 150B，共约 70 层（整体范围 2–4T 参数，70–90 层）。Sol 极可能采用轻量 KV cache 设计（类似 DeepSeekV4 或混合 SSM），以充分利用 SRAM 带宽。这表明 OpenAI 并非简单将前沿模型部署到第三方平台，而是围绕硬件设计了该模型，使 Cerebras 从“小型模型快速推理”跃升至服务前沿智能的极致速度。

## 正文

OpenAI is launching GPT-5.6 Sol on Cerebras at up to 750 tokens per second in July.

Bleys Goodson estimates Sol may be served across 70-100 Cerebras wafers, with roughly one model layer per wafer: around 3T total parameters, 150B active, 70 layers.

It suggests OpenAI did not just place a frontier model onto another inference provider. It likely designed the model around the hardware!
Cerebras would move from “fast inference for smaller models” into something much more strategic: serving frontier intelligence at extreme speed.

Game changer, big win for OpenAI. It cannot be emphasized enough how important this is.

### 引用推文

> Bleys Goodson：It is a 2 to 4T param model. They are serving it across 70-100 wafers. To get healthy serving characteristics, they are essentially putting at most one layer pe...
