随着人工智能多年来的发展,其带来的影响也在逐渐加剧。模型的能力越来越强,我们的工作方式正在迅速改变,AI 的经济效益正变得真实可感,与此同时,现实世界的风险也日益凸显。2026 年将是我认为这种势头不会出现任何停歇的第一年。最难准备的部分在于,情况很有可能从此持续升级——更多的颠覆、更多的意外、更高的赌注。
就我而言,有一系列话题对我如何看待当前 AI 的现状至关重要,但我甚至还没来得及写下来(至少没有从所有我想要的角度去写!)。所有这些都紧密关联着不同模型达到新能力水平所带来的影响,以及我如何据此推断接下来可能发生的事情。
1. 开源模型尚未迎来像 Opus 4.5 那样的真正智能体时刻
开源与闭源模型之间的时间差距经常被讨论,但现实是,我们有一个很好的时间窗口,它独立于有争议的基准测试——即开源权重模型是否能在智能体框架中变得极其有用。2025 年 12 月 Claude Code 中 Opus 4.5 的时刻是如此响亮和明显,以至于如果开源模型能以低至每月 5 美元的价格达到这一性能水平,使用量将会出现爆炸式增长。
目前我们已经过去了大约 5 到 6 个月,还没有出现同等级别的开源模型。我怀疑,我所撰写的那些最佳闭源前沿模型的稳健性,可能会让这个时刻需要更长的时间才能到来,比如说接近 12 个月以上。在这段时间里,Claude Code 和 Codex 可能会看起来像是不同类别的产品。在来自不同实验室的新一代最先进开源模型的常规热潮中,基准测试分数肯定会继续攀升,但随着实际使用成为真正的试金石,开源与闭源之间的差距应该会变得更加清晰可辨。
2. Gemini 仍然没有出现能与 Claude Code 和 Codex 相抗衡的有力竞争者
我能为“开源模型实际落后程度比基准测试显示的更严重”这一预测提供的最有力佐证是:就连强大的谷歌,目前也没有一款能明确与 Claude Code 和 Codex 抗衡的产品。我相信 Gemini 团队正在这方面全力推进。
我仍需要对 Gemini 3.5 Flash 进行更多测试,但阅读评测后可以清楚地看到,它无法替代我目前的工作方式。也许并非 Gemini 团队刻意针对谷歌现有产品(搜索、YouTube 等)进行专门优化,但该模型似乎确实适合这些场景。如果谷歌短期内拿不出强大的工具,我也不指望开源模型实验室能做到。开源模型将更多地用于自动化企业级智能体以及低成本领域,而非成为现代知识工作的核心驱动工具。这将直接影响未来模型的资金引擎——像 Claude Code 和 Codex 这样的智能体,正是当前实现 AI 收入大规模增长的最佳路径。
我曾与 Grace Shao 讨论过,当前环境正悄然推动中国实验室专注于 AI 领域(Proem),而这正是我预测未来几年开源模型将走向专业化、而非与 OpenAI、Anthropic 和谷歌直接竞争的核心依据。
3. 我不认为今年会出现开源权重的 Mythos
虽然我不认为 Mythos 是一个能在所有领域碾压对手的通用“神级模型”,但我确实认为它在软件工程和网络安全领域是一项卓越的技术成就。Mythos 显然是这些领域的里程碑事件。在与大多数中国实验室——尤其是那些拥有最突出、大规模开源 MoE 模型的机构,如 Kimi、Z.ai、DeepSeek 和通义千问——交流后,我认为它们资源严重受限,短期内无法像美国大型实验室那样扩大训练规模。而对于那些更具企业背景、资源也更充足的实验室,如阿里巴巴和字节跳动,它们在安全与保障方面则持更为保守的立场。Mythos 是美国顶级公司在训练和研究算力上实现巨大加速的风向标。
Epoch AI 最近发布了一篇关于各实验室可用算力的精彩文章(谷歌约 25%,Meta 11%,OpenAI 11%,Anthropic 6%)。所有这些数字都远高于任何中国实验室。
4. 美国开源模型正逐渐蓄力
英伟达的 Nemotron、谷歌的 Gemma、Arcee AI 等公司正在逐步稳定美国开源模型生态系统。这方面有很多难以衡量的因素,尤其是像 OpenClaw 和 Hermes 这类本地智能体的兴起,但美国模型的使用量数据达到了自 Llama 3 以来我们从未见过的水平。Gemma 4 的各款模型在性能上均与同等规模的 Qwen 3.5/3.6 模型持平或更优——而长期以来,Qwen 一直是该规模下默认的开源模型。这些 Qwen 3.5/3.6 模型在许多后训练研究中难以顺利使用,部分原因在于架构/工具链,部分原因可能在于建模本身(即该模型因某些训练决策而不易微调)。我很少听到对 Gemma 的抱怨,但这也可能是因为 Gemma 尚未成为研究人员的默认选择。
我们最近从 GPT-OSS、Nemotron 3 以及现在的 Gemma 4 等模型中看到了一个简单的事实:如果一个模型在基准测试中处于合适的水平,并且由美国实验室以真正宽松的许可证发布,那么它将会获得大量采用(回想一下,在这一轮中,Gemma 4 采用了 Apache 2.0 许可证,改变了早期 Gemma 模型带有使用限制的许可方式)。美国在开源模型领域的这一早期增长阶段,正在直接与开发者建立关键品牌认知。共识是,像 Reflection 和 Thinking Machines 这样的新实验室很可能会参与这一领域,但过于耐心将会错失构建新的智能体工作流和企业关系的关键时机。
5. Anthropic 和 OpenAI 在模型迭代上才刚刚提速
我预计今年剩余时间将是这两家旗舰公司之间的残酷竞争。我目前处于一个有趣的平衡点:我认为 GPT 5.5 是更聪明的模型,而且我很喜欢 Codex 应用,因此我正在将大部分工作安排成能在该平台上完成。与此同时,对于许多与写作相关以及更广泛领域的任务,我仍然非常喜欢 Claude。这些模型正在迅速改变我们的工作方式——我会在忙其他事情时用手机运行 Codex,正在基于智能体设置自动化开源模型分析任务,并且期望能够大规模扩展 Interconnects 的研究侧。
人工智能正开始推动公司走向规模化时代的两个极端。最大的公司将比以往任何时候都庞大,利用资源和海量人才在原始 AI 能力的前沿取得持续进步。另一方面,像 Interconnects 这样的小型公司则通过使用智能体来提炼、展示和销售细分领域的专业知识而蓬勃发展。即将到来的大规模社会性岗位替代,将降低那些在纯技术层面(无论大公司还是小公司)都不属于这两个极端的各类知识工作者的就业能力,同时维持甚至可能放大那些直接与人类打交道(例如医生)或拥有其他能够自我维持的权力结构(法律/政府)的职业。
6. 更多现有权力结构将在 AI 领域彰显自身力量
就在撰写本文的过去几天里,教皇发布了一份超过四万字的关于 AI 发展方向的文档,而中国则扩大了对顶尖 AI 研究人员在行业间流动的限制。与此同时,美国已将 Anthropic 列为供应链风险,并继续将其模型用于国家安全。这类新闻只会越来越多。现有权力结构正在意识到,它们能在 AI 动态中施加影响力的时间窗口是有限的——这种直觉可以理解为,随着 AI 模型变得更强大,其影响力会下降。这种直觉可能很危险,因为它会在谁控制这项技术的问题上引发重大冲突(正如我在 Anthropic 与 DoW 争执后与 Dean Ball 讨论的那样)。
接下来:技术问题如何演变为社会问题
这些大体上属于技术层面和权力格局的加速趋势,将给美国国内日益高涨的反 AI 社会与政治情绪带来更大压力。这目前是 AI 持续发展和有益扩散最明显的障碍。反思这一点,科技讨论中的许多人过于关注细节,诚然,许多数据中心反对者确实在为其立场辩护时提出了完全错误的事实性主张。
相当一部分美国人真正的立场是:他们有权对当前趋势说“不”——通过拒绝批准建设数据中心。这是科技行业在过去几十年里改变了全球经济格局和权力结构后,从未赋予他们的发言权。
这正为行业未来一年带来严峻挑战。各大实验室正在将人才聚集和集中到前所未有的水平。几乎没有中立的信使能向公众传达 AI 的真实状况。前沿实验室的领导层大多正筹备 IPO,并力求在能力竞赛中保持领先。在现状下,几乎没有什么行动能扭转这条通往社会冲突的道路。
这需要 AI 生态系统中的个体另辟蹊径,逆流而行,摆脱那种“今天就必须发财”、“必须在实验室才能做有影响力的工作”等群体思维。我个人继续押注于此,努力打造一个由清晰、无偏见信息支持的、充满活力且多元化的开放模型生态系统。如果你认同这一点,并且一直在旁观,那么现在是时候参与进来了——在局势失控之前。
As the years of AI progress go by, it’s been accompanied by a slowly rising tide of consequence. Models are getting more capable, how we work is changing quickly, economics of AI are becoming real, just as real-world risks come to the forefront. 2026 is the first year where I don’t think there’ll be any breaks from this. The hard part to prepare for is that there’s a good chance things just continue to ratchet up from here – more disruption, more surprises, more stakes.
On my end, there’s been a growing list of topics that are very fateful to how I see the current state of AI, but I haven’t even gotten to write about them (at least not from all the angles I want to)! All of these are closely related to the implications of different models reaching new capability levels and how I use that to infer what may come next.
1. Open models haven’t had their true agent moment like Opus 4.5
The time gap between open and closed models is very often discussed, but the reality is that we have a nice time-gating that’s independent of debatable benchmarks – if open-weight models do or do not become super useful in agentic harnesses. The Opus 4.5 in Claude Code moment of December 2025 was so loud and obvious, that if open models hit this performance level for price points as low as $5/month, there will be an explosion in usage.
Right now we are about 5-6 months in with no equivalent open model. I suspect the robustness of the best closed frontier models that I write about could make this moment take a good amount longer, say closer to 12+ months. In this time, Claude Code and Codex may seem like different categories of products. In the standard flurry of new, state-of-the-art open models from a variety of labs, benchmarks will definitely keep climbing, but the open-closed gap should become more interpretable as real-world use becomes the real litmus test.
2. Gemini still doesn’t have a meaningful competitor for Claude Code and Codex
The best exclamation point I can offer to reinforce my prediction that open models are further behind than the benchmarks claim is that even the mighty Google doesn’t have a clear competitor for Claude Code and Codex. I’m sure the Gemini team is pushing very hard on this.
I still need to do a lot more testing on Gemini 3.5 Flash, but reading reviews makes it clear that it’s not a substitute for how I’m working today. It’s maybe not the Gemini team explicitly specializing for Google’s existing products (search, YouTube, etc.), but the model seems to suit them. If Google doesn’t have a powerful tool here soon, I don’t expect the open model labs to either. The open models are going to be used more for automated, enterprise agents and low-cost domains, rather than being the driving tool of modern knowledge work. This will feed directly into the economic engine of funding future models, where the agents like Claude Code and Codex are the current best path to massive AI revenue growth.
I discussed how the current environment is quietly driving labs in China to specialize on AI Proem with Grace Shao and this is central to my expectations of open models specializing over the next few years instead of competing with OpenAI, Anthropic, and Google.
3. I don’t expect an open-weights Mythos this year
While I don’t think Mythos is a general “god model” that will crush the competition in every domain, I do think it’s a remarkable technical achievement in software engineering and cybersecurity. Mythos is obviously a watershed moment for those fields. Having spoken to most of the Chinese labs – particularly those with the most prominent, large, open MoE models like Kimi, Z.ai, DeepSeek, and Qwen – I think they’re heavily resource limited and don’t have an immediate path to scaling up training processes like the big labs in the U.S. For the labs which are more corporate, which comes with more resources, such as Alibaba and Bytedance, they also have more conservative stances on safety and security.
Mythos is a bellwether of the massive acceleration in training and research compute available to the largest American companies.
Epoch AI recently had a nice piece on the compute available to various labs (~Google 25%, Meta 11%, OpenAI 11%, Anthropic 6%). All of these numbers are vastly higher than any Chinese lab.
4. American open models are slowly gaining steam
Nvidia with Nemotron, Google with Gemma, Arcee AI and others are slowly stabilizing the open model ecosystem in the U.S. There’s a lot that’s hard to measure here, especially in the rise of local agents like OpenClaw and Hermes, but there are adoption numbers of American models that we haven’t seen since Llama 3.
Gemma 4’s models are all tying or outperforming the equivalently sized Qwen 3.5/3.6 models — where Qwen has for years now been the default open model at these sizes. These Qwen 3.5/3.6 models have been tricky to get working in a lot of post-training research, partially due to architecture/tooling and partially likely due to modeling (i.e. the model is not easy to finetune for some training decision). I’ve heard few complaints about Gemma, but it also could be because Gemma is not yet the researcher default.
There's a simple reality that we've seen recently with models like GPT-OSS, Nemotron 3, and now Gemma 4, that if a model is in the right range of benchmarks and released by an American lab with a truly permissive license, it'll get a large amount of adoption (in this cycle, recall that Gemma 4 adopted the Apache 2.0 License, changing from one with use-case restrictions on earlier Gemmas). This early phase of American growth in open models is establishing key brands directly with developers. The consensus is that more neolabs like Reflection and Thinking Machines are likely to participate in this space, but being too patient will lose the time when new agentic workflows and enterprise relationships are built.
5. Anthropic and OpenAI are just getting up to speed in model iterations
I expect the rest of this year to be a ruthless competition between these two flagship companies. I’m at an interesting balance where I think GPT 5.5 is a bit smarter of a model and I love the Codex App, so I’m structuring much of my work to be possible there. At the same time, for a lot of writing-related and broader surface area tasks I really still love Claude. These models are rapidly changing how we work, I run Codex from my phone while doing other things, am setting up automated open model analysis jobs on the back of agents, and expect to be able to scale the research side of Interconnects widely.
AI is beginning to drive companies to the two extremes in the scaling era. The biggest companies will be way bigger than ever, using resources and mass talent to have sustained progress at the frontier of raw AI capabilities. On the other side, tiny businesses like Interconnects thrive by using agents to refine, present, and sell niche expertise. The mass social job displacement that’ll come is going to reduce employability for various knowledge workers that don’t fit into either of these extremes for the raw technical side (big or small companies), while sustaining and maybe even amplifying careers that interface directly with humans (e.g. doctors) or other power structures with means to sustain themselves (law/government).
6. More existing power structures will assert themselves on AI
Just in the last few days while writing this, we had the Pope release an over 40,000 word document on where AI is goingand China expand personnel movement restrictions on top AI researchers across industry. At the same time, the U.S. has designated Anthropic a supply chain risk and continues to use its models for national security. The list of news like this is only going to grow. Existing power structures are realizing there’s a finite time window for them to exert themselves in the AI dynamic — an intuition that could be mapped to influence going down as AI models get more powerful. This intuition is potentially dangerous, as it sets up meaningful conflict in who controls the technology (as I discussed with Dean Ball after the Anthropic-DoW spat).
Next: Where technical becomes social
These largely technical and power trends accelerating are going to put more pressure on the social and political anti-AI sentiments within the U.S. This is currently the most obvious barrier to continued AI development and beneficial diffusion. Reflecting on this, many people in the tech discourse get too focused on the details, where yes a lot of data-center-detractors are making genuinely wrong factual claims in defense of their position.
The real position that a large swath of Americans has is that they have a voice in saying no to the current trend — by not granting permission to build data centers. This is a voice that they haven’t been granted by the tech industry that changed the face of the global economy and power structures in the last few decades.
This is setting us up for a challenging year ahead for the industry. The labs are aggregating and concentrating talent to peak levels. There are few neutral messengers to communicate the reality of AI to the public. The frontier labs leadership is largely gearing up to IPO and stay ahead in the capabilities race. With the status quo, there are few actions to unwind this path toward social conflict.
It takes individuals in the AI ecosystem to zag and go against the groupthink of needing to make your wealth today, of needing to be at a lab to do impactful work, and so on. I’m personally continuing to bet on this, by trying to make a vibrant and diverse open model ecosystem supported by clear, unbiased information. If you agree with this and have been watching from the sidelines, it’s a good time to get involved, before the situation spirals into something uncontrollable.