Cerebras CS-4 发布:推理性能翻倍

Chubby♨️ · @kimmonismus · X·2026-08-20 23:57·14天前
AI 导读

Cerebras 发布 CS-4,未换新制程,仅通过重新设计供电与散热将时钟速度翻倍,使 AI 推理性能近乎翻番。单晶圆提供 250 PFLOPs 算力、43.2 PB/s 内存带宽;单机柜集成三片晶圆,达 750 PFLOPs。在 GPT-OSS-120B 上,每用户每秒超 4,400 tokens,推理速度较 GPU 系统最高快 30 倍。

Chubby♨️@kimmonismus
41AI 编辑部评分,满分 100

Cerebras CS-4 发布:推理性能翻倍

2026-08-20 23:57· 14天前
AI 导读

Cerebras 发布 CS-4,未换新制程,仅通过重新设计供电与散热将时钟速度翻倍,使 AI 推理性能近乎翻番。单晶圆提供 250 PFLOPs 算力、43.2 PB/s 内存带宽;单机柜集成三片晶圆,达 750 PFLOPs。在 GPT-OSS-120B 上,每用户每秒超 4,400 tokens,推理速度较 GPU 系统最高快 30 倍。

This is seriously impressive: Cerebras just unveiled CS-4, and nearly doubled its AI inference performance without moving to a new process node.

Same gigantic 5nm wafer. Same 4 trillion transistors. Same 900,000 AI cores. Instead, Cerebras redesigned the power delivery and cooling, allowing the wafer to run at twice the clock speed. The result per WSE-3 Turbo:

• 250 PFLOPs of AI compute • 43.2 PB/s of memory bandwidth • 2.4 Tb/s of I/O bandwidth

A single CS-4 rack combines three wafers for 750 PFLOPs and 129.6 PB/s of memory bandwidth.

On GPT-OSS-120B, Cerebras reports more than 4,400 tokens per second per user, and up to 30x faster inference than GPU-based systems.

So yeah, intelligence not only too cheap to meter but also too fast to keep up with