The Decoder:AI News(RSS)
54AI 编辑部评分,满分 100

GPT-5.6 Sol 推出 Ultrafast 模式,推理速度提升 14 倍,由 Cerebras 提供支持

2026-08-14 22:21· 8分钟前· Matthias Bastian
AI 导读

OpenAI 预览 GPT-5.6 Sol 的 Ultrafast 模式,输出速度最高达每秒 750 token,较此前提升 14 倍,加速由 Cerebras 提供支持,双方今年早些时候签署了百亿美元合作协议。该服务初期仅通过 OpenAI API 向特定客户开放,OpenAI 计划随算力增长逐步扩大访问。OpenAI 已在内部将该模式用于事件响应,并计划将速度作为新的定价杠杆。

Image description

OpenAI is launching a preview of its new "Ultrafast" mode, which delivers up to 750 output tokens per second from its flagship model, GPT-5.6 Sol.

The inference acceleration comes from Cerebras, which signed a ten-billion-dollar partnership with OpenAI earlier this year. The service will initially be available only through the OpenAI API for GPT-5.6 Sol and limited to select customers. OpenAI plans to expand access gradually as capacity grows. Companies that want in can sign up for updates through a form.

Faster output opens up new use cases

OpenAI says Ultrafast is designed to combine the speed of smaller models with the full capabilities of a large reasoning model, enabling what the company calls "more useful work per second." During incident response, for example, engineers could have logs, code changes, and reports analyzed while an outage is still happening, helping them pinpoint the cause and prepare a fix in real time. OpenAI says it's already using the model internally for this purpose.

OpenAI pitches several other scenarios. In finance, the model could evaluate market signals and flag suspicious transactions while conditions are still shifting. In customer support, complex inquiries could be resolved in real time, even when finding the answer requires multiple steps or systems. In e-commerce, it could answer product questions, check inventory levels, and personalize recommendations before a hesitant buyer abandons their cart.

视频 · 前往原文观看

Video: Generating a 3D warehouse simulator, shown on the left in Ultrafast mode

Research is another area OpenAI highlights. Experiments that previously ran overnight as batch jobs could turn into interactive work sessions, letting teams test an idea, review results, adjust their approach, and kick off another run without breaking their workflow.

OpenAI turns AI speed into a pricing lever

OpenAI already monetizes inference speed in tiers. Through the API, it offers a "Fast Mode" that promises up to 2.5x speed with lower latency for GPT-5.6 Sol at roughly double the price. Ultrafast adds a third, faster, and likely pricier tier.

The logic mirrors cloud providers like AWS, which have long charged more for the same service at higher performance levels. OpenAI is applying that to AI inference. If speed becomes a bottleneck across industries, this tiered model gives OpenAI a direct cut of the revenue gains that faster inference creates.

来源:The Decoder:AI News(RSS) · the-decoder.com

GPT-5.6 Sol 推出 Ultrafast 模式,推理速度提升 14 倍,由 Cerebras 提供支持

The Decoder:AI News(RSS)·2026-08-14 22:21·8分钟前·Matthias Bastian
AI 导读

OpenAI 预览 GPT-5.6 Sol 的 Ultrafast 模式,输出速度最高达每秒 750 token,较此前提升 14 倍,加速由 Cerebras 提供支持,双方今年早些时候签署了百亿美元合作协议。该服务初期仅通过 OpenAI API 向特定客户开放,OpenAI 计划随算力增长逐步扩大访问。OpenAI 已在内部将该模式用于事件响应,并计划将速度作为新的定价杠杆。

原文 · 保持原样,未翻译
Image description

OpenAI is launching a preview of its new "Ultrafast" mode, which delivers up to 750 output tokens per second from its flagship model, GPT-5.6 Sol.

The inference acceleration comes from Cerebras, which signed a ten-billion-dollar partnership with OpenAI earlier this year. The service will initially be available only through the OpenAI API for GPT-5.6 Sol and limited to select customers. OpenAI plans to expand access gradually as capacity grows. Companies that want in can sign up for updates through a form.

Faster output opens up new use cases

OpenAI says Ultrafast is designed to combine the speed of smaller models with the full capabilities of a large reasoning model, enabling what the company calls "more useful work per second." During incident response, for example, engineers could have logs, code changes, and reports analyzed while an outage is still happening, helping them pinpoint the cause and prepare a fix in real time. OpenAI says it's already using the model internally for this purpose.

OpenAI pitches several other scenarios. In finance, the model could evaluate market signals and flag suspicious transactions while conditions are still shifting. In customer support, complex inquiries could be resolved in real time, even when finding the answer requires multiple steps or systems. In e-commerce, it could answer product questions, check inventory levels, and personalize recommendations before a hesitant buyer abandons their cart.

视频 · 前往原文观看

Video: Generating a 3D warehouse simulator, shown on the left in Ultrafast mode

Research is another area OpenAI highlights. Experiments that previously ran overnight as batch jobs could turn into interactive work sessions, letting teams test an idea, review results, adjust their approach, and kick off another run without breaking their workflow.

OpenAI turns AI speed into a pricing lever

OpenAI already monetizes inference speed in tiers. Through the API, it offers a "Fast Mode" that promises up to 2.5x speed with lower latency for GPT-5.6 Sol at roughly double the price. Ultrafast adds a third, faster, and likely pricier tier.

The logic mirrors cloud providers like AWS, which have long charged more for the same service at higher performance levels. OpenAI is applying that to AI inference. If speed becomes a bottleneck across industries, this tiered model gives OpenAI a direct cut of the revenue gains that faster inference creates.

来源:The Decoder:AI News(RSS)· the-decoder.com