SemiAnalysis 揭秘 OpenAI Jalapeño 芯片

Rohan Paul · @rohanpaul_ai · X·2026-08-26 07:27·38分钟前
AI 导读

SemiAnalysis 长文披露 OpenAI 自研 Jalapeño 芯片,称其面向数十万亿参数模型与百万级 token 上下文设计,并质疑 CUDA 在 AI 可自写代码优化新架构下的存续。该芯片以每瓦性能为核心,目标让数千芯片协同为单一推理机,其每兆瓦输出 token 吞吐已超越英伟达 Vera Rubin 公布数据。

Rohan Paul@rohanpaul_ai
44AI 编辑部评分,满分 100

SemiAnalysis 揭秘 OpenAI Jalapeño 芯片

2026-08-26 07:27· 38分钟前
AI 导读

SemiAnalysis 长文披露 OpenAI 自研 Jalapeño 芯片,称其面向数十万亿参数模型与百万级 token 上下文设计,并质疑 CUDA 在 AI 可自写代码优化新架构下的存续。该芯片以每瓦性能为核心,目标让数千芯片协同为单一推理机,其每兆瓦输出 token 吞吐已超越英伟达 Vera Rubin 公布数据。

SemiAnalysis wrote a super long piece on OpenAI's Jalapeño chip.

Some revelations

• OpenAI may be designing infrastructure for: models measured in tens of trillions of parameters or context windows containing millions of tokens.

• They questioned if CUDA can survive when AI itself can rapidly write and optimize software for a completely new architecture.

• The chip showed for the first time that chips may no longer need perfect universal compilers if frontier models can write the hard parts themselves.

"If Jalapeño is a success, it will be a strong signal that the industry’s obsession over programming models and perfect, universal compilers are invalidated by frontier AI models."

• “OpenAI designs for perf/W.” OpenAI appears to be optimizing around a different scarce resource: not money, and not floor space, but electricity. Once power becomes the hard ceiling, tokens per watt starts looking like the real currency of AI infrastructure.

• Jalapeño isn't being designed as an isolated accelerator. OpenAI is building the networking architecture to make thousands of its own chips behave like one enormous inference machine.

• “Jalapeño smokes every other chip.” Specifically about token throughput per unit of datacenter power. In an industry increasingly constrained by electricity, that may be a more important victory than raw benchmark speed.

• Blackwell is almost the easier comparison. The much more provocative claim is that Jalapeño can already beat published results from Nvidia’s newer Vera Rubin generation on output-token throughput per megawatt.

来源:Rohan Paul· x.com