OpenAI 将在 Cerebras 上推出 GPT-5.6 Sol,推理速度达 750 tokens/s

Chubby♨️ · @kimmonismus · X·2026-07-06 15:39·57天前
AI 导读

OpenAI 将于 7 月在 Cerebras 上推出 GPT-5.6 Sol,推理速度达 750 tokens/s。据 Bleys Goodson 估算,模型横跨 70–100 片 Cerebras 晶圆,每片约承载一层,总参数量约 3T,活跃参数 150B,共约 70 层(整体范围 2–4T 参数,70–90 层)。Sol 极可能采用轻量 KV cache 设计(类似 DeepSeekV4 或混合 SSM),以充分利用 SRAM 带宽。这表明 OpenAI 并非简单将前沿模型部署到第三方平台,而是围绕硬件设计了该模型,使 Cerebras 从“小型模型快速推理”跃升至服务前沿智能的极致速度。

Chubby♨️@kimmonismus
73AI 编辑部评分,满分 100

OpenAI 将在 Cerebras 上推出 GPT-5.6 Sol,推理速度达 750 tokens/s

2026-07-06 15:39· 57天前
AI 导读

OpenAI 将于 7 月在 Cerebras 上推出 GPT-5.6 Sol,推理速度达 750 tokens/s。据 Bleys Goodson 估算,模型横跨 70–100 片 Cerebras 晶圆,每片约承载一层,总参数量约 3T,活跃参数 150B,共约 70 层(整体范围 2–4T 参数,70–90 层)。Sol 极可能采用轻量 KV cache 设计(类似 DeepSeekV4 或混合 SSM),以充分利用 SRAM 带宽。这表明 OpenAI 并非简单将前沿模型部署到第三方平台,而是围绕硬件设计了该模型,使 Cerebras 从“小型模型快速推理”跃升至服务前沿智能的极致速度。

OpenAI is launching GPT-5.6 Sol on Cerebras at up to 750 tokens per second in July.

Bleys Goodson estimates Sol may be served across 70-100 Cerebras wafers, with roughly one model layer per wafer: around 3T total parameters, 150B active, 70 layers.

It suggests OpenAI did not just place a frontier model onto another inference provider. It likely designed the model around the hardware! Cerebras would move from “fast inference for smaller models” into something much more strategic: serving frontier intelligence at extreme speed.

Game changer, big win for OpenAI. It cannot be emphasized enough how important this is.

Bleys GoodsonIt is a 2 to 4T param model. They are serving it across 70-100 wafers. To get healthy serving characteristics, they are essentially putting at most one layer pe...