Only two labs occupy the Time per Task Pareto frontier: all frontier models under two minutes are from OpenAI, and four of the six above are from Anthropic
We show the tradeoff between the Artificial Analysis Intelligence Index and how long it takes models to complete a representative task in the Index. We calculate Time per Task as the number of output tokens divided by output (decode) speed. Contributors to Time per Task include inference optimizations such as speculative decoding, and model architecture factors (number of active parameters, amount of reasoning).
All the models on the frontier under two minutes per task are from OpenAI - but above this, four of the six are Anthropic models. Kimi K3 (max) has a notably high ~11 minutes per task, driven by a low output speed on Moonshot AI's first party API.
Note, the chart only shows models scoring above 25 on the Artificial Analysis Intelligence Index