OpenAI is launching GPT-5.6 Sol on Cerebras at up to 750 tokens per second in July.
Bleys Goodson estimates Sol may be served across 70-100 Cerebras wafers, with roughly one model layer per wafer: around 3T total parameters, 150B active, 70 layers.
It suggests OpenAI did not just place a frontier model onto another inference provider. It likely designed the model around the hardware! Cerebras would move from “fast inference for smaller models” into something much more strategic: serving frontier intelligence at extreme speed.
Game changer, big win for OpenAI. It cannot be emphasized enough how important this is.