This is seriously impressive: Cerebras just unveiled CS-4, and nearly doubled its AI inference performance without moving to a new process node.
Same gigantic 5nm wafer. Same 4 trillion transistors. Same 900,000 AI cores. Instead, Cerebras redesigned the power delivery and cooling, allowing the wafer to run at twice the clock speed. The result per WSE-3 Turbo:
• 250 PFLOPs of AI compute • 43.2 PB/s of memory bandwidth • 2.4 Tb/s of I/O bandwidth
A single CS-4 rack combines three wafers for 750 PFLOPs and 129.6 PB/s of memory bandwidth.
On GPT-OSS-120B, Cerebras reports more than 4,400 tokens per second per user, and up to 30x faster inference than GPU-based systems.
So yeah, intelligence not only too cheap to meter but also too fast to keep up with