就像马拉松比赛中的两艘帆船,开放与封闭的 AI 实验室正在旧金山湾中迎风转向、灵活穿梭。
2023 年,闭源模型在 Chatbot Arena Elo 评分上以巨大优势领先。两年后,DeepSeek R1 时刻到来,这是对 ChatGPT 时刻的开源回应。两艘船并驾齐驱了近一年。
架构改进以及首批基于 Blackwell 训练的模型,从 2026 年开始带来了 GPT-5.2 和 Fable 5 的阶跃式变化。
未来几周将迎来又一轮开源发布潮。月之暗面于 7 月 16 日发布了 Kimi K3,这是一个 2.8T 参数的开源权重模型。阿里巴巴于 7 月 19 日预览了通义千问 3.8,这是一个 2.4T 参数的模型。DeepSeek V4 于 7 月中旬从预览版转正。紧随其后的是 Thinking Machines 的 Inkling(7 月 15 日发布的 975B Apache-2.0 多模态模型),以及 Meta Superintelligence Labs 于 4 月发布的 Muse Spark。
开源模型从未在公开水域取得领先,但这或许并无必要。按 90/10 的输入输出价格比混合计算,中位水平的开源前沿模型运行成本比 GPT-5.2 便宜约 15%。最便宜的开源模型 DeepSeek V4 Flash 则便宜约 90%。
我们可能会看到这样一种格局:闭源模型推动行业向前发展,而开源模型迅速跟进并将其商品化。这会减缓创新吗?
竞争往往会产生相反的效果。OpenAI 已将推理成本降低了 50%。Kimi 推出了新的注意力架构 KDA。Fable 的阶跃式进步正促使整个行业加倍努力迎头赶上。
摆在行业面前的主要问题是利润率将走向何方。Anthropic 即将公布其首个盈利季度。贝索斯曾说,你的利润就是我的机会。开源的竞争格局使得利润率和定价始终保持竞争力。
AI 浪潮将成为美国有史以来最大的基础设施项目之一,也很可能是推动经济更快增长的最大贡献者之一。竞争对于保持这场竞赛的速度至关重要。
前沿竞赛不再是一场单向的赛跑。它是一个循环往复的过程:闭源模型领先,开源模型追赶,整个市场随之加速前进。
-
Chatbot Arena Elo 是一种源自国际象棋的评分体系。用户会并排看到两个匿名模型的回答,并投票选出更好的一个。每个模型初始分为 1000 分。战胜更强的模型比战胜更弱的模型获得更多分数;评分差距可以预测某场对局的胜率。100 分的 Elo 差距意味着评分更高的模型大约有 64% 的胜率。该分数反映的是人类在开放式对话中的偏好,而非推理、编程或智能体相关基准的表现。↩︎
-
关于规模估算及其传导的增长渠道,请参阅《大语言模型的 GDP 影响》。↩︎
Like two sailboats in a marathon race, open & closed AI labs are tacking & jibing in San Francisco Bay.
In 2023, closed source models led by an enormous margin on Chatbot Arena Elo1. Two years later, the DeepSeek R1 moment arrived, the open-source answer to the ChatGPT moment. The two boats raced side-by-side for nearly a year.
Architectural improvements & the first Blackwell-trained models brought a step change with GPT-5.2 & Fable 5 starting in 2026.
The next few weeks will see another flurry of open-source releases. Moonshot shipped Kimi K3, a 2.8T parameter open-weight model, on July 16. Alibaba previewed Qwen 3.8, a 2.4T model, on July 19. DeepSeek V4 graduates from preview in mid-July. These follow Thinking Machines’ Inkling, a 975B Apache-2.0 multimodal model released July 15, & Meta Superintelligence Labs’ Muse Spark in April.
Open-source models have never taken an open-water lead, but that may not be necessary. Blend prices at a 90/10 input-to-output ratio & the median open-weight frontier model runs about 15% cheaper than GPT-5.2. The cheapest open model, DeepSeek V4 Flash, is roughly 90% cheaper.
We may have a dynamic where the closed models drive the industry forward & open-source rapidly copies to commoditize. Will that slow down innovation?
Competition tends to do the opposite. OpenAI has cut inference costs by 50%. Kimi shipped a new attention architecture, KDA. Fable’s step function has an entire industry redoubling to catch up.
The major question put to the industry is what will happen to margins. Anthropic is about to post its first profitable quarter. Bezos said your margin is my opportunity. Open source’s competitive dynamics keep margins & pricing competitive.
The AI wave will be among the largest infrastructure projects2 ever for the US & likely one of the greatest contributors to faster economic growth. Competition is essential to keeping the race fast.
The frontier is no longer a one-way race. It is a repeating cycle: closed models pull ahead, open models catch up, & the whole market moves faster.
-
Chatbot Arena Elo is a rating system borrowed from chess. Users see responses from two anonymous models side-by-side & vote for the better one. Each model starts at 1000. Winning against a stronger model earns more points than winning against a weaker one; the gap in ratings predicts the probability of winning a matchup. A 100-point Elo gap implies the higher-rated model wins about 64% of the time. The score reflects human preference on open-ended chat, not reasoning, coding, or agentic benchmarks. ↩︎
-
See The GDP Impact of LLMs for the scale estimate & the growth channel it flows through. ↩︎