最先进的模型比去年十一月时聪明了三分之二。这种飞速改进的节奏仍在持续,每三天就有两款新模型问世。
但在 OpenRouter 上,84% 的 token 并非来自最先进的模型。
事实上,用户选择用来生成绝大多数 token 的六款模型,其性能约为前沿模型的 77%,而成本仅为 Claude Fable 5 的 2.5%。
该指数持续攀升。大约每个季度就会出现一次三到五个 Artificial Analysis 积分的大幅跃升,较小的涨幅则填补了其间空白。
在 8 月 10 日那一周,六款模型承载了 80% 的流量。它们的综合价格为每百万 token 0.50 美元,而 Fable 5 为 20 美元。
Ramp 的数据显示,买家对价格高度敏感。Fable 5 定价约为每百万 token 10 美元,上线一个月后便占据了 Anthropic token 流量的 6% 和 Anthropic 支出的 11%。GPT-5.6 Sol 作为 OpenAI 最贵的主流产品线,则持有 OpenAI 约四分之一的 token 流量。
尽管 Fable 5 价格远高于 GPT-5.6 Sol,但七月份其模型归因收入约为后者的 75%。
每一代新的最先进模型发布,所抢占的市场份额都应少于前一代。
企业将整合支出。合同会像云时代那样集中在少数一两家供应商手中,而一旦某款模型通过了高价值任务的考验,该工作负载就会长期留在那里。
在显著的价格折扣下,性能已经足够出色。差距正从下方不断收窄。最好的开源权重模型到五月份时已达到前沿模型分数的 80%,而一年前这一比例仅为 48%。
应用部署则是另一回事。我们越来越多的投资组合公司和初创企业默认选择更小的模型、微调模型和开源模型。他们在优化另一条帕累托前沿——以价格换性能。
如果市场份额不再转移,且“够用就好”持续成立,那么最先进模型的经济逻辑就会改变。一次九位数的训练投入必须赢得市场份额才能收回成本,而这个门槛会随着时间推移不断抬高。
这个标题有些轻率。很多人确实在购买最先进的模型,而且理由充分。软件工程架构和安全性设计是最明显的例子,在这些场景中,最好的可用模型物有所值。
但我们掌握的公开数据表明,真正重要的前沿是另一条。
-
Artificial Analysis 模型目录与智能指数。主要实验室的月度发布数量与前沿路径。样本起始于 2025-11-01。发布速率趋势持平。大型(≥3 分)前沿跃迁之间的中位间隔约为 3.5 个月。智能指数 ↩︎
-
最先进水平指某一周内可获得的最佳 Artificial Analysis 得分;当某模型的得分处于该周最佳命名模型的 10% 以内时,即被视为接近前沿。将 OpenRouter 每周命名的头部模型与 Artificial Analysis 得分进行匹配,取的是 OpenRouter 轮播列表的头部,而非全部 API。份额序列,周次为 2025-11-03 至 2026-05-25(n=30),仅计命名模型,排除 Others。前十三周与后十三周相比,接近前沿的比例约为 17.5% 对 14.6%(约 82-85% 处于前沿之外)。集中度快照,2026-08-10 当周,覆盖命名 token 前约 80% 的模型。按 token 加权的 Artificial Analysis 得分落后全球目录最先进水平约 23%(约为前沿质量的 77%),并落后该 OpenRouter 列表上最佳模型约 10%。混合篮子约为 $0.50/百万 token,而 Fable 5 为 $20/百万(约 40 倍)。2026 年 5 月的历史核查显示,落后本地列表领先者约 16%。最佳开源权重模型在前十三周的前沿得分为 47.5%,后十三周为 70.9%;单周最佳为 2026-05-25,DeepSeek V4 Pro 得分 45.27,前沿为 56.31(80.4%)。OpenRouter 排名 ↩︎ ↩︎ ↩︎
-
这些数据源并未涵盖第一方云服务,即 OpenAI、Anthropic 与 Google 自身的服务。运行在原生 API 上的前沿流量从未进入 OpenRouter 排名,因此数据存在偏差。 ↩︎
-
Ramp Economics Lab,AI 指数 2026 年 8 月(Fable 5 采用情况)。econlab.substack.com/p/ai-index-august-2026 ↩︎
State of the art models are two-thirds smarter than they were last November. The frenetic pace of improvement is sustained, two new models every three days. 1
But 84% of tokens on OpenRouter aren’t state of the art. 2 3
In fact, the six models users choose to generate the supermajority of those tokens deliver about 77% of the performance of the frontier. They cost 2.5% of what Claude Fable 5 does. 2
The index keeps jumping. Large gains of three to five Artificial Analysis points land about every quarter. Smaller steps fill the gaps.
Six models carry 80% of volume in the week of August 10. Their blended price is $0.50 per million tokens against Fable 5 at $20.
Ramp’s data shows buyers are price-elastic. Fable 5 at about $10/m tokens captured 6% of Anthropic tokens & 11% of Anthropic spend a month after launch. GPT-5.6 Sol, OpenAI’s priciest mainline tier, held about a quarter of OpenAI tokens. 4
Fable 5 generated roughly 75% as much model-attributed revenue as GPT-5.6 Sol in July, despite being substantially more expensive.
Each new state of the art release should move less share than the one before it.
Enterprises will consolidate spend. Contracts concentrate on one or two vendors, just like in the cloud era, & once a model clears a high-value job the workload stays.
Performance is already good enough at a meaningful discount. The gap keeps closing from below. The best open-weight model reached 80% of the frontier score by May, up from 48% a year earlier. 2
Application deployment is the other story. More of our portfolio companies & startups default to smaller models, fine-tuned models, & open source. They are optimizing against a different Pareto frontier, price over performance.
If share stops shifting & good enough stays good enough, the economics of SOTA change. A nine-figure training run has to win share to pay for itself, & that bar will rise with time.
The title is flippant. Plenty buy SOTA, & for good reason. Software engineering architecture & security design are the clearest cases, where the best available model earns its price.
But the open data we do have suggests the frontier that matters is the other one.
-
Artificial Analysis model catalog & Intelligence Index. Major-lab monthly release counts & frontier path. Sample starts 2025-11-01. Release-rate trend flat. Median gap between large (≥3 pt) frontier steps about 3.5 months. Intelligence Index ↩︎
-
State of the art means the single best Artificial Analysis score available in a given week; a model counts as near it when the score sits within 10% of that week’s best named model. OpenRouter weekly named top models joined to Artificial Analysis scores, the head of the OpenRouter carousel rather than every API. Share series, weeks 2025-11-03 through 2026-05-25 (n=30), named only, Others excluded. First vs last thirteen weeks about 17.5% vs 14.6% near the frontier (~82-85% outside). Concentration snapshot, week of 2026-08-10, models covering the first ~80% of named tokens. Token-weighted Artificial Analysis about 23% behind global catalog state of the art (~77% of frontier quality) & about 10% behind the best model on that OpenRouter list. Blended basket about $0.50/m tokens vs Fable 5 at $20/m (~40x). May 2026 historical check, about 16% behind the local list leader. Best open-weight model 47.5% of frontier score in the first thirteen weeks vs 70.9% in the last thirteen; single best week 2026-05-25, DeepSeek V4 Pro at 45.27 vs frontier 56.31 (80.4%). OpenRouter rankings ↩︎ ↩︎ ↩︎
-
These data sources don’t capture the first-party clouds, OpenAI, Anthropic & Google’s own services. Frontier traffic running on native APIs never enters the OpenRouter rankings, so there’s a bias to the data. ↩︎
-
Ramp Economics Lab, AI Index August 2026 (Fable 5 uptake). econlab.substack.com/p/ai-index-august-2026 ↩︎