Semianalysis on OpenAI's new Jalapeño chips that will be deployed inside its own compute by the end of 2026.
"Jalapeño smokes every other chip. All this is done without Multi Token Prediction (MTP), while the other chips on the chart are the best performing configs of each respective SKU, all with MTP"
Jalapeño shifts the entire latency–efficiency Pareto frontier upward: at ~100 tok/s/user it delivers ~11M tok/s/MW, roughly 2× the best Blackwell-class configs at comparable interactivity.
The kicker: that’s STP (Single-Token Prediction) with no speculative decoding, while the competing curves are already using MTP.