Tomer Tunguz 博客(VC 分析)
54AI 编辑部评分,满分 100

谁真的需要SOTA模型?OpenRouter数据显示84%token来自非前沿模型

2026-08-14 08:00· 1天前
AI 导读

OpenRouter数据显示,84%的模型token并非来自SOTA模型,用户最常用的六款模型性能约为前沿模型的77%,成本仅为Claude Fable 5的2.5%。8月10日当周,六款模型承载了80%流量,混合价格约$0.50/百万token,而Fable 5为$20。最佳开源模型性能已从一年前的48%提升至前沿模型的80%,企业正转向更小、微调或开源模型以优化性价比。

State of the art models are two-thirds smarter than they were last November. The frenetic pace of improvement is sustained, two new models every three days. 1

Monthly major-lab model releases since November 2025, averaging about 20 per month

But 84% of tokens on OpenRouter aren’t state of the art. 2 3

Non state of the art token share on OpenRouter named top models held near 84 percent

In fact, the six models users choose to generate the supermajority of those tokens deliver about 77% of the performance of the frontier. They cost 2.5% of what Claude Fable 5 does. 2

The index keeps jumping. Large gains of three to five Artificial Analysis points land about every quarter. Smaller steps fill the gaps.

Step gains when a new model sets the Artificial Analysis intelligence frontier

Six models carry 80% of volume in the week of August 10. Their blended price is $0.50 per million tokens against Fable 5 at $20.

OpenRouter top models at 77% of state of the art quality for one-fortieth the Fable price

Ramp’s data shows buyers are price-elastic. Fable 5 at about $10/m tokens captured 6% of Anthropic tokens & 11% of Anthropic spend a month after launch. GPT-5.6 Sol, OpenAI’s priciest mainline tier, held about a quarter of OpenAI tokens. 4

Fable 5 generated roughly 75% as much model-attributed revenue as GPT-5.6 Sol in July, despite being substantially more expensive.

Each new state of the art release should move less share than the one before it.

Enterprises will consolidate spend. Contracts concentrate on one or two vendors, just like in the cloud era, & once a model clears a high-value job the workload stays.

Performance is already good enough at a meaningful discount. The gap keeps closing from below. The best open-weight model reached 80% of the frontier score by May, up from 48% a year earlier. 2

Application deployment is the other story. More of our portfolio companies & startups default to smaller models, fine-tuned models, & open source. They are optimizing against a different Pareto frontier, price over performance.

If share stops shifting & good enough stays good enough, the economics of SOTA change. A nine-figure training run has to win share to pay for itself, & that bar will rise with time.

The title is flippant. Plenty buy SOTA, & for good reason. Software engineering architecture & security design are the clearest cases, where the best available model earns its price.

But the open data we do have suggests the frontier that matters is the other one.


  1. Artificial Analysis model catalog & Intelligence Index. Major-lab monthly release counts & frontier path. Sample starts 2025-11-01. Release-rate trend flat. Median gap between large (≥3 pt) frontier steps about 3.5 months. Intelligence Index ↩︎

  2. State of the art means the single best Artificial Analysis score available in a given week; a model counts as near it when the score sits within 10% of that week’s best named model. OpenRouter weekly named top models joined to Artificial Analysis scores, the head of the OpenRouter carousel rather than every API. Share series, weeks 2025-11-03 through 2026-05-25 (n=30), named only, Others excluded. First vs last thirteen weeks about 17.5% vs 14.6% near the frontier (~82-85% outside). Concentration snapshot, week of 2026-08-10, models covering the first ~80% of named tokens. Token-weighted Artificial Analysis about 23% behind global catalog state of the art (~77% of frontier quality) & about 10% behind the best model on that OpenRouter list. Blended basket about $0.50/m tokens vs Fable 5 at $20/m (~40x). May 2026 historical check, about 16% behind the local list leader. Best open-weight model 47.5% of frontier score in the first thirteen weeks vs 70.9% in the last thirteen; single best week 2026-05-25, DeepSeek V4 Pro at 45.27 vs frontier 56.31 (80.4%). OpenRouter rankings ↩︎ ↩︎ ↩︎

  3. These data sources don’t capture the first-party clouds, OpenAI, Anthropic & Google’s own services. Frontier traffic running on native APIs never enters the OpenRouter rankings, so there’s a bias to the data. ↩︎

  4. Ramp Economics Lab, AI Index August 2026 (Fable 5 uptake). econlab.substack.com/p/ai-index-august-2026 ↩︎

来源:Tomer Tunguz 博客(VC 分析) · tomtunguz.com

谁真的需要SOTA模型?OpenRouter数据显示84%token来自非前沿模型

Tomer Tunguz 博客(VC 分析)·2026-08-14 08:00·1天前
AI 导读

OpenRouter数据显示,84%的模型token并非来自SOTA模型,用户最常用的六款模型性能约为前沿模型的77%,成本仅为Claude Fable 5的2.5%。8月10日当周,六款模型承载了80%流量,混合价格约$0.50/百万token,而Fable 5为$20。最佳开源模型性能已从一年前的48%提升至前沿模型的80%,企业正转向更小、微调或开源模型以优化性价比。

原文 · 保持原样,未翻译

State of the art models are two-thirds smarter than they were last November. The frenetic pace of improvement is sustained, two new models every three days. 1

Monthly major-lab model releases since November 2025, averaging about 20 per month

But 84% of tokens on OpenRouter aren’t state of the art. 2 3

Non state of the art token share on OpenRouter named top models held near 84 percent

In fact, the six models users choose to generate the supermajority of those tokens deliver about 77% of the performance of the frontier. They cost 2.5% of what Claude Fable 5 does. 2

The index keeps jumping. Large gains of three to five Artificial Analysis points land about every quarter. Smaller steps fill the gaps.

Step gains when a new model sets the Artificial Analysis intelligence frontier

Six models carry 80% of volume in the week of August 10. Their blended price is $0.50 per million tokens against Fable 5 at $20.

OpenRouter top models at 77% of state of the art quality for one-fortieth the Fable price

Ramp’s data shows buyers are price-elastic. Fable 5 at about $10/m tokens captured 6% of Anthropic tokens & 11% of Anthropic spend a month after launch. GPT-5.6 Sol, OpenAI’s priciest mainline tier, held about a quarter of OpenAI tokens. 4

Fable 5 generated roughly 75% as much model-attributed revenue as GPT-5.6 Sol in July, despite being substantially more expensive.

Each new state of the art release should move less share than the one before it.

Enterprises will consolidate spend. Contracts concentrate on one or two vendors, just like in the cloud era, & once a model clears a high-value job the workload stays.

Performance is already good enough at a meaningful discount. The gap keeps closing from below. The best open-weight model reached 80% of the frontier score by May, up from 48% a year earlier. 2

Application deployment is the other story. More of our portfolio companies & startups default to smaller models, fine-tuned models, & open source. They are optimizing against a different Pareto frontier, price over performance.

If share stops shifting & good enough stays good enough, the economics of SOTA change. A nine-figure training run has to win share to pay for itself, & that bar will rise with time.

The title is flippant. Plenty buy SOTA, & for good reason. Software engineering architecture & security design are the clearest cases, where the best available model earns its price.

But the open data we do have suggests the frontier that matters is the other one.


  1. Artificial Analysis model catalog & Intelligence Index. Major-lab monthly release counts & frontier path. Sample starts 2025-11-01. Release-rate trend flat. Median gap between large (≥3 pt) frontier steps about 3.5 months. Intelligence Index ↩︎

  2. State of the art means the single best Artificial Analysis score available in a given week; a model counts as near it when the score sits within 10% of that week’s best named model. OpenRouter weekly named top models joined to Artificial Analysis scores, the head of the OpenRouter carousel rather than every API. Share series, weeks 2025-11-03 through 2026-05-25 (n=30), named only, Others excluded. First vs last thirteen weeks about 17.5% vs 14.6% near the frontier (~82-85% outside). Concentration snapshot, week of 2026-08-10, models covering the first ~80% of named tokens. Token-weighted Artificial Analysis about 23% behind global catalog state of the art (~77% of frontier quality) & about 10% behind the best model on that OpenRouter list. Blended basket about $0.50/m tokens vs Fable 5 at $20/m (~40x). May 2026 historical check, about 16% behind the local list leader. Best open-weight model 47.5% of frontier score in the first thirteen weeks vs 70.9% in the last thirteen; single best week 2026-05-25, DeepSeek V4 Pro at 45.27 vs frontier 56.31 (80.4%). OpenRouter rankings ↩︎ ↩︎ ↩︎

  3. These data sources don’t capture the first-party clouds, OpenAI, Anthropic & Google’s own services. Frontier traffic running on native APIs never enters the OpenRouter rankings, so there’s a bias to the data. ↩︎

  4. Ramp Economics Lab, AI Index August 2026 (Fable 5 uptake). econlab.substack.com/p/ai-index-august-2026 ↩︎

来源:Tomer Tunguz 博客(VC 分析)· tomtunguz.com