Artificial Analysis@ArtificialAnlys
53AI 编辑部评分,满分 100
2026-08-06 03:20· 34分钟前
AI 导读

Artificial Analysis 显示,任务耗时帕累托前沿仅由两家实验室占据:2 分钟内的前沿模型全部来自 OpenAI,而 2 分钟以上的 6 个模型中 4 个来自 Anthropic。Kimi K3(max)因 Moonshot AI 首方 API 输出速度低,单任务耗时约 11 分钟。该图仅展示智能指数超 25 分的模型。

Only two labs occupy the Time per Task Pareto frontier: all frontier models under two minutes are from OpenAI, and four of the six above are from Anthropic

We show the tradeoff between the Artificial Analysis Intelligence Index and how long it takes models to complete a representative task in the Index. We calculate Time per Task as the number of output tokens divided by output (decode) speed. Contributors to Time per Task include inference optimizations such as speculative decoding, and model architecture factors (number of active parameters, amount of reasoning).

All the models on the frontier under two minutes per task are from OpenAI - but above this, four of the six are Anthropic models. Kimi K3 (max) has a notably high ~11 minutes per task, driven by a low output speed on Moonshot AI's first party API.

Note, the chart only shows models scoring above 25 on the Artificial Analysis Intelligence Index

来源:Artificial Analysis · x.com

Artificial Analysis · @ArtificialAnlys · X·2026-08-06 03:20·34分钟前
AI 导读

Artificial Analysis 显示,任务耗时帕累托前沿仅由两家实验室占据:2 分钟内的前沿模型全部来自 OpenAI,而 2 分钟以上的 6 个模型中 4 个来自 Anthropic。Kimi K3(max)因 Moonshot AI 首方 API 输出速度低,单任务耗时约 11 分钟。该图仅展示智能指数超 25 分的模型。

Only two labs occupy the Time per Task Pareto frontier: all frontier models under two minutes are from OpenAI, and four of the six above are from Anthropic

We show the tradeoff between the Artificial Analysis Intelligence Index and how long it takes models to complete a representative task in the Index. We calculate Time per Task as the number of output tokens divided by output (decode) speed. Contributors to Time per Task include inference optimizations such as speculative decoding, and model architecture factors (number of active parameters, amount of reasoning).

All the models on the frontier under two minutes per task are from OpenAI - but above this, four of the six are Anthropic models. Kimi K3 (max) has a notably high ~11 minutes per task, driven by a low output speed on Moonshot AI's first party API.

Note, the chart only shows models scoring above 25 on the Artificial Analysis Intelligence Index

来源:Artificial Analysis· x.com