坐在从杭州开往上海的新型高速列车上,我凝视窗外,层峦叠嶂的山脊线点缀着风力发电机,在夕阳的映照下形成剪影,美不胜收。群山为广袤的田野和密集的摩天大楼群提供了壮丽的背景。我带着极大的谦逊从中国归来。去到一个如此陌生的地方,却受到如此热情的欢迎,这是一次非常温暖、充满人情味的经历。我有幸见到了许多在AI生态系统中素未谋面的人,他们用灿烂的笑容和欢呼迎接我,让我意识到我的工作以及整个AI生态系统是多么具有全球性。
中国研究者的心态
构建语言模型的中国公司,其定位是这项技术的完美快速追随者,这建立在长期以来的教育和文化传统之上,同时也采用了略有不同的科技公司构建方式。当你审视其产出(支持智能体工作流的最新、最大模型)以及其要素(优秀的科学家、大规模数据和加速计算)时,中国和美国的实验室看起来大体相似。而持久的差异则体现在这些要素的组织和条件设定方式上。
我长期以来一直认为,中国实验室如此擅长追赶并保持在前沿水平的一个原因是,他们在文化上就适合这项任务。但如果没有直接与人交流,我觉得自己无权将这种直觉归因于实质性的影响。与来自中国顶尖实验室的许多优秀、谦逊且开放的科学家交谈后,我的许多信念变得更加清晰。
如今,构建最优秀的大语言模型,很大程度上取决于贯穿整个技术栈的细致工作,从数据到架构细节,再到强化学习算法的实现。模型的每个环节都能带来一些改进,而将它们整合在一起是一个复杂的过程,在此过程中,一些杰出个人的工作可能需要被搁置,以服务于整个模型在多目标优化中实现最大化。
美国研究人员显然也擅长解决各个独立组件的问题,但在美国,更盛行一种为自己发声的文化。作为一名科学家,当你为自己的工作发声时,你会更成功,而现代文化正在推动“顶尖AI科学家”这条新的成名之路。这导致了直接的冲突。据广泛传闻,Llama 组织已经在这种将个人利益嵌入层级组织的政治压力下崩溃了。我听说其他实验室表示,有时需要付钱给一位顶尖研究员,才能让他们停止抱怨自己的想法没有被纳入最终模型。无论这是否完全属实,其核心思想是明确的。自我意识和职业晋升的欲望确实会阻碍打造最佳模型。中美之间这种文化上的微小方向性差异,会对最终产出产生重大影响。
这其中部分原因与在中国构建模型的人员构成有关。所有实验室都面临一个现实情况,即核心贡献者中有很大一部分是在校学生。这些实验室非常年轻,这让我想起了我们在 Ai2 的架构,在那里学生被视为同侪,并直接融入大语言模型团队。这与美国顶尖实验室截然不同,像 OpenAI、Anthropic、Cursor 等公司根本不提供实习机会。其他公司如谷歌,名义上有与 Gemini 相关的实习项目,但很多人担心你的实习会被孤立起来,接触不到任何实质性的工作。
总结一下,这种文化上的细微变化如何能够提升构建模型的能力:
更愿意为了改进最终模型而从事不那么光鲜亮丽的工作,
刚进入 AI 领域的人可以不受之前 AI 炒作周期的影响,从而能够更快地适应新的现代技术(事实上,我交谈过的一位中国科学家非常积极地认同这一优势),
较少的自我意识使得组织架构能够略微扩展,因为系统内的博弈行为减少了,以及
充足的人才非常适合解决那些在其他地方已有概念验证的问题,等等。
这种略微偏向于补充构建当今语言模型所需技能的倾向,与一个众所周知的刻板印象形成对比,即中国研究者往往产出较少的创造性、开创领域、从0到1的学术型研究。在我们此行访问的更具学术氛围的实验室中,许多领导者都在谈论培养这种更具雄心的研究文化。与此同时,我们交谈过的一些技术领导者对此持怀疑态度,认为短期内不太可能实现这种科研方式的重新调整,因为这需要重新设计教育和激励机制,其规模之大在当前的经济均衡状态下难以实现。这种文化似乎正在培养出在构建大语言模型方面极为出色的学生和工程师。当然,他们的人数也极其充裕。
这些学生告诉我,中国也出现了与美国类似的人才流失现象,许多原本考虑走学术道路的人现在打算留在产业界。最有趣的一句话来自一位研究者,他原本有兴趣成为教授以接近教育体系,但评论说教育问题已经被大语言模型解决了——“学生为什么要跟我交流!”
这些学生有一个优势,他们能以全新的视角看待大语言模型。过去几年里,我们看到大语言模型的关键范式从扩展MoE,转向扩展强化学习,再到实现智能体。要做好其中任何一项,都需要快速吸收海量的上下文信息,既包括广泛的文献,也包括所在公司的技术栈。学生们习惯于这样做,并且乐于谦逊地放下所有关于什么应该有效的预设。他们一头扎进去,全身心投入,只为获得改进模型的机会。
这些学生还出奇地直率,摆脱了那些可能分散科学家注意力的哲学空谈。当被问及他们对模型的经济影响或长期社会风险有何看法时,很少有中国研究者持有复杂的观点或想要影响这些方面。他们的角色就是构建最好的模型。
这种差异很微妙,也容易被否认,但当你与一位优雅、才华横溢、能用英语清晰交流的研究员进行长时间对话时,最能感受到这一点。关于人工智能更哲学层面的基本问题,会因简单的困惑而悬在空中。对他们来说,这是一种范畴错误。一位研究员甚至在探讨这些领域时引用了著名的 Dan Wang 论断——中国由工程师管理,而美国由律师管理——以强调他们想要构建的愿望。在中国,没有一条系统性地培养中国科学家明星力量的路径,类似于 Dwarkesh 或 Lex 这类大型主流播客。
试图让中国科学家评论由人工智能引发的即将到来的经济不确定性、超越简单通用人工智能能力的问题,或关于模型应如何行为的道德辩论,都充分体现了这些科学家极度的谦逊。这不仅仅是专注于工作,而是他们不想对自己不了解的问题发表评论。
放眼全局——北京尤其感觉像湾区,一个竞争激烈的实验室只需步行或打车就能到达。我下了飞机,在去酒店的路上顺道去了阿里巴巴的北京园区。然后,在 36 小时内,我们走访了 Z.ai、月之暗面、清华大学、美团、小米和零一万物。用滴滴出行很方便,如果你在中国选择 XL 车型,通常会遇到带有按摩座椅的电动小型货车。我们向研究员们询问了人才争夺战,他们表示这与我们在美国经历的情况非常相似。研究员跳槽很常见,人们选择去哪很大程度上取决于当前哪里的氛围最好。
在中国,大语言模型社区感觉更像一个生态系统,而不是相互争斗的部落。在许多非公开的对话中,只有对同行的尊重。所有中国实验室都害怕字节跳动及其流行的豆包模型,这是中国唯一一家前沿闭源实验室。同时,所有实验室都极度尊重深度求索,认为它是执行层面研究品味最好的实验室。当你与美国实验室成员进行非公开会面时,火花会很快迸发。
中国研究人员谦逊态度中最引人注目的地方在于,他们在商业层面也常常耸耸肩,表示那不是他们的问题;而在美国,似乎每个人都痴迷于各种生态系统层面的产业趋势,从数据销售商到算力或融资,无一不关注。
中国 AI 产业与西方实验室的差异(与趋同之处)
如今构建 AI 模型之所以如此有趣,是因为它不再仅仅是把一群优秀的研究人员聚集在同一栋楼里,打造出一个工程奇迹。过去或许如此,但要维持 AI 业务,大语言模型正逐渐成为构建、部署、融资以及推动这项创造被采纳的混合体。领先的 AI 公司存在于复杂的生态系统之中,这些系统提供资金、算力、数据等资源,以持续推动前沿发展。
对于创造和维持大语言模型所需的各种要素的整合,西方生态系统(以 Anthropic 和 OpenAI 为代表)已经有了相当清晰的概念化和图谱化。因此,发现中国实验室在思考方式上的重大差异,就能看出不同公司可能对未来做出截然不同的押注。当然,这些未来在很大程度上会受到融资和/或算力方面的制约。
我记录了与这些实验室交流后,在“AI 产业”层面得出的最大收获:
国内 AI 需求的早期迹象。有一种广为流传的假设认为,中国 AI 市场会更小,因为中国公司往往不愿为软件付费——因此无法释放出支撑实验室发展的巨大推理市场。这种说法仅适用于映射到 SaaS 生态系统的软件支出,而 SaaS 在中国历来规模很小;另一方面,中国显然仍然存在一个庞大的云市场。一个关键且悬而未决的问题——中国实验室内部也在争论——是企业对 AI 的支出究竟会跟随 SaaS 市场(规模小)还是云市场(基础性)。总体来看,感觉 AI 正更趋近于云市场,而且没有人对新工具周围正在形成的市场感到特别担忧。
大多数开发者都是 Claude 的忠实拥趸。中国绝大多数 AI 开发者都对 Claude 及其对软件开发方式的改变痴迷不已,尽管 Claude 名义上在中国是被禁止的。仅仅因为中国历来对购买软件持谨慎态度,并不能让我认为推理需求不会出现爆发式增长。中国的技术人员非常务实、谦逊且充满动力——这一事实似乎比过去不花钱的习惯性承诺更为强大。一些中国研究人员提到会使用自己的工具进行开发,比如 Kimi 或 GLM 的命令行界面,但所有人都提到会使用 Claude 进行开发。此外,提到 Codex 的次数也出奇地少,而它在湾区显然正越来越受欢迎。
中国企业有一种技术所有权心态。中国文化正与强劲的经济引擎相结合,催生出难以预测的结果。我始终有一种感觉,那就是众多 AI 模型反映了这里许多科技企业务实且当下的平衡状态。并没有什么宏大的总体规划。这个行业的特点是对字节跳动和阿里巴巴的尊重,这些行业巨头凭借其雄厚资源,有望在几乎所有市场中占据很大份额。深度求索是备受尊敬的技术领导者,但远非市场领导者。它们设定了方向,但并未做好在经济上获胜的准备。这就给美团或蚂蚁集团这样的公司留下了空间,西方人可能会惊讶于它们也在构建这些模型。实际上,它们清楚地看到大语言模型显然是未来科技产品的核心,因此需要一个强大的基础。当它们对强大且通用的模型进行微调时,这既加固了自身的技术栈,又能从开源社区获得反馈,同时还能为自家产品保留内部的微调版本。行业中“开放优先”的心态很大程度上是由实用性定义的——这有助于让它们的模型获得强有力的反馈,回馈开源社区,并赋能其使命。
政府援助确实存在,但规模尚不明确。人们常声称中国政府正积极助力开源大语言模型竞赛。这是一个权力分散于多级政府的体系,每一级政府对于自身具体职责并无清晰行动指南。北京的各区会竞相吸引科技公司将办公地点设在其辖区内。这些公司获得的“帮助”几乎肯定包括简化行政审批等官僚程序,但这种帮助究竟能走多远?各级政府能否协助吸引人才?能否帮助走私芯片?在整个访问期间,多次提及政府的兴趣或帮助,但细节信息太少,不足以断言其具体内容,也无法形成关于政府如何影响中国人工智能发展轨迹的明确世界观。当然,没有任何迹象表明中国政府高层在影响模型的技术决策。
数据行业的发展远未成熟。我们听闻过 Anthropic 或 OpenAI 为单一环境投入超过 1000 万美元,累计每年花费数亿美元来推动强化学习的前沿,因此我们迫切想知道中国实验室是直接从美国公司购买相同的环境,还是由国内镜像生态系统提供支持。答案并非完全不存在数据行业,而是他们的经验表明,数据行业质量相对较差,通常最好在内部自行构建环境或数据。研究人员自身会花费大量时间制作强化学习训练环境,而字节跳动、阿里巴巴等一些大公司则拥有内部数据标注团队来支持这项工作。这一切都印证了上一点中提到的“自建而非购买”的思路。
对更多英伟达芯片的迫切需求。英伟达的计算资源是训练的黄金标准,每家公司的进展都因无法获得更多此类资源而受限。如果供应充足,他们显然会购买。其他加速器,包括但不限于华为的产品,在推理方面获得了积极评价。无数实验室都已获得华为芯片的使用权限。
这些观点描绘了一幅截然不同的人工智能生态图景:若试图将西方实验室的运作模式快速套用到中国同行身上,往往会导致范畴错误。关键问题在于,这些不同的生态系统是否会产生本质不同的模型类型,抑或中国模型始终会被解释为与美国3至9个月前的前沿模型相似。
结论:全球均衡
我深知自己在启程前对中国知之甚少,而此行结束时,只觉刚刚开始学习。中国并非能用规则或公式概括的地方,而是一个拥有截然不同的动态机制与内在化学反应的国度。这里的文化如此古老、如此深厚,且至今仍与本土技术的构建方式完全交织在一起。我面前还有很长的学习之路。
美国当前权力结构中的许多决策,都严重依赖其对中国现有的世界观作为关键思维工具。在与几乎所有中国领先的人工智能实验室进行过正式或非正式的面对面交流后,我发现中国有许多特质和本能,是西方决策模式极难模拟的。即便直接追问这些实验室为何公开其顶级模型,所有权意识与对生态系统的真诚支持之间的交集,仍让我难以理清其中的关联。
这里的实验室非常务实,并非开源领域的绝对主义者——并非它们构建的每一个模型都会公开——但在支持开发者、滋养生态系统,并将其作为更深入了解自身模型的手段方面,存在着深层的意图性。
几乎所有中国大型科技公司都在构建自己的通用大语言模型,正如我们看到美团(外卖服务平台)和小米(综合性消费科技公司)等企业发布了开源权重模型。美国同类公司通常只会购买相关服务。这些公司开发大语言模型并非为了追逐热点以保持存在感,而是出于一种深层的根本性渴望——掌控自身技术栈,并开发当今最重要的技术。当我从笔记本电脑前抬起头,总能看到地平线上成排的塔吊,这显然与更广泛的中国建设文化及其蕴含的能量相契合。
中国研究者身上的人性光辉、人格魅力与真诚热情,极具人性化感染力。在个人层面,美国习以为常的那种你死我活的地缘政治话语,在他们身上完全看不到。这个世界需要更多这样纯粹的积极态度。作为AI社群的一员,我目前更担忧的是,围绕国籍标签在成员和群体之间出现的裂痕。
如果说我不希望美国实验室在AI技术栈的每个层面都保持明确领先地位——尤其是在我深耕的开源模型领域——那是在撒谎。我是美国人,这是诚实的偏好。与此同时,我希望开源生态系统本身能在全球蓬勃发展,因为这能为世界创造更安全、更易获取、更有用的AI。而当前的问题是,美国实验室是否会采取行动来占据这一领导地位。
在完成本文之际,更多关于行政命令影响开源模型的传闻正在发酵,这可能会进一步复杂化美国领导力与全球生态系统之间的协同关系——这并未让我感到乐观。
感谢我在月之暗面、智谱、美团、小米、通义千问、蚂蚁集团、零一万物及其他公司有幸交流的所有杰出人士。每个人都如此热情好客,慷慨地付出时间。随着我对中国的认知逐渐清晰,我将继续分享我的思考,涵盖广义的文化层面以及具体的AI领域。显然,这些知识将直接关系到AI发展前沿正在展开的叙事。
Staring out the window on a new, high-speed train from Hangzhou to Shanghai I’m gifted with views of dramatic ridgelines speckled with wind turbines that are silhouetted against the setting sun. The mountains cast a backdrop to a mix of spanning fields and clustered skyscrapers. I’m returning from China with great humility. It’s a very warming, human experience to go somewhere so foreign and be so welcomed. I had the honor of meeting so many people in the AI ecosystem who I knew from afar, and they greeted me with big smiles and cheer, reminding me how global my work and the AI ecosystem is.
The mentality of Chinese researchers
The Chinese companies building language models are set up as the perfect fast-followers for the technology, building on long-standing cultural traditions in education and work, along with subtly different approaches to building technology companies. When you look at the outputs, the latest, biggest models enabling agentic workflows, and the ingredients, excellent scientists, large-scale data, and accelerated computing, the Chinese and American labs look largely similar. The lasting differences emerge in how these are organized and conditioned.
I’ve long thought that a reason that the Chinese labs are so good at catching up and keeping up with the frontier is that they’re culturally aligned for this task, but without talking to people directly I felt like it wasn’t my place to attribute substantial influence to this hunch. Speaking with many wonderful, humble, and open scientists at the leading Chinese labs has crystallized a lot of my beliefs.
So much of building the best LLMs today comes down to meticulous work across the entire stack, from data to architecture details and RL algorithm implementations. All points of the model can give some improvements, and fitting them in together is a complex process where the work of some brilliant individuals needs to get shelved in favor of the overall model maximizing a multi-objective optimization.
Where American researchers are obviously also brilliant at solving the individual components, there’s more of a culture of speaking up for yourself in the U.S. As a scientist, you’re more successful when you speak up for your work and modern culture is pushing the new path to fame of “leading AI scientists”. This results in direct conflict. The Llama organization is heavily rumored to have collapsed under the political weight of these interests embedding themselves in a hierarchical organization. I’ve heard of other labs saying that it can be needed to pay off a top researcher to get them to stop complaining about their idea not making it in the final model. Whether or not that’s exactly true, the idea is clear. Ego and desires for career advancement do get in the way of making the best models. A small, directional shift in this sort of culture between the U.S. and China can have a meaningful impact on the final outputs.
Some of this has to do with who is building the models in China. There’s an immediate reality at all of the labs that a large proportion of the core contributors are active students. The labs are quite young, and it reminds me of our setup at Ai2, where students are seen as peers and directly integrated in the LLM team. This is incredibly different from the top labs in the US, where the likes of OpenAI, Anthropic, Cursor, etc. simply don’t offer internships. Other companies like Google nominally have internships related to Gemini, but there’s a lot of concern about whether your internship will be siloed and away from anything real.
To summarize how the slight change in culture can improve the ability to build models:
More willingness to do non-flashy work in order to improve the final model,
People new to building AI can be free of prior phases of AI hype cycles, allowing them to adapt to the new modern techniques faster (in fact, one of the Chinese scientists I talked to really actively attached to this strength),
Less ego enabling org charts to scale slightly, as there’s less gamifying the system, and
Abundant talent well-suited to solving problems with a proof of concept elsewhere, etc.
This slight inclination towards skills that complement building today’s language models stands in contrast to a known stereotype that Chinese researchers tend to produce less creative, field-spawning, 0-to-1 academic style research. Among the more academic lab visits on our trip, many leaders talk about cultivating this more ambitious research culture. At the same time, some technical leaders we talked to were skeptical about whether such a rewiring in the approach to science is likely in the near term, because it’ll take a redesign of the education and incentive systems that is too big to happen within the current economic equilibrium. This culture seems to be training students and engineers that are excellent at the LLM building game. They also, of course, have an extremely abundant quantity.
These students told me about a similar brain drain happening in China as in the U.S., where many who previously considered academic paths now intend to stay in industry. The funniest quote was from a researcher who was interested in being a professor to be close to the education system, but remarked that education is solved with LLMs – “why would a student talk to me!”
The students have a benefit of coming at LLMs with fresh eyes. Over the last few years we’ve seen the key paradigm of LLMs shift from scaling MoE’s, to scaling RL, to enabling agents. Doing any of these well involves absorbing an insane amount of context quickly, both from the broader literature and the technical stack at your company. Students are used to doing this and excited to humbly drop all presumptions about what should work. They dive in head first and dedicate their life to getting the chance to improve the models.
These students are also so magically direct and free of some of the philosophical chatter that can distract scientists. When asking questions on how they feel about the economics or long-term social risks of models, far fewer Chinese researchers have sophisticated opinions and a drive to influence this. Their role is to build the best model.
This difference is subtle, and easy to deny, but it is best felt when having long conversations with an elegant, brilliant researcher who can clearly communicate well in English, basic questions on more philosophical aspects of AI hang in the air with a simple confusion. It’s a category error to them. One researcher even quoted the famous Dan Wang premise of China being run by engineers, relative to the lawyers of the U.S. when probing in these areas, to emphasize their desire to build. There’s no track in China that systematically enables the growth of star power for Chinese scientists, akin to mega mainstream podcasts like Dwarkesh or Lex.
Trying to get Chinese scientists to comment on the coming economic uncertainty fueled by AI, questions beyond the capabilities of simple AGI, or moral debates on how models should behave all served to capture the extreme humility of these scientists. It’s more than just being dedicated to their work, but they don’t want to comment on issues they’re not informed on.
Zooming out — Beijing especially felt much like the Bay Area, where a competitive lab is a short walk or Uber away. I got off a flight and stopped by Alibaba’s Beijing campus on the way to the hotel. Then, in 36 hours we went to all of Z.ai, Moonshot AI, Tsinghua University, Meituan, Xiaomi, and 01.ai. Travel by Didi is easy, and if you select an XL in China you’re often paired with electric mini vans that have massage chairs. We asked the researchers about the talent wars, and they said it’s very similar to what we’re experiencing in the U.S. It’s normal for researchers to bounce around, and much of where people choose to go is based on the best current vibes.
In China, the LLM community feels far more like an ecosystem than battling tribes. Across many off the record conversations, it’s nothing but respect for peers. All of the Chinese labs fear Bytedance with their popular Doubao model, which is the only frontier closed lab in China. At the same time, all of the labs have massive respect for DeepSeek as the lab with the best research taste in execution. When you meet with lab members off the record in the States, sparks fly quickly.
The most striking part of the humility of Chinese researchers is how they also often shrug on the business side, saying it’s not their problem, where everyone in the U.S. seems to be obsessed with various ecosystem-level industrial trends, from data sellers to compute or fundraising.
Where China’s AI industry differs (and matches) the Western labs
The thing that makes building an AI model today so interesting is that it’s not just about getting a group of great researchers in one building together to produce an engineering marvel. It used to be this, but to sustain AI businesses, the LLMs are becoming a mix of building, deploying, funding, and getting adoption for this creation. The leading AI companies exist in complex ecosystems that supply money, compute, data and more in order to keep pushing the frontier.
The integration of these various inputs to creating and sustaining LLMs is fairly well conceptualized and mapped for the Western ecosystem, as typified by Anthropic and OpenAI, so finding big differences in how the Chinese labs think about it points at where the different companies can be making meaningfully different bets on the future. Of course, these futures can be heavily dictated by the constraints on funding and/or compute.
I’ve documented the biggest “AI Industry” level take-aways from talking to these labs:
Early signs of domestic AI demand. There’s a much-touted hypothesis that the Chinese AI market will be smaller because Chinese companies don’t tend to pay for software – thus, never unlocking a giant inference market supporting labs. This is only true for software spend that maps to the SaaS ecosystem, which is historically tiny in China, where on the other hand there is obviously still a large cloud market in China. A crucial unanswered question – one which the Chinese labs themselves debate – on if spending for AI in the enterprise tracks the SaaS market (small) or the cloud market (fundamental). On net, it feels like AI is trending closer to the cloud, and no one was actively worried about a market growing around the new tools.
Most developers are Claude-pilled. Most of the AI developers in China are obsessed with Claude and how it’s changed how they build software, despite Claude nominally being banned in China. Just because China has historically been hesitant to buy software does not give me the impression that there won’t be a massive surge in inference demand. Chinese technical staff are so practical, humble, and motivated – a fact that seems stronger than any commitment to previous habits in not spending.
Some Chinese researchers mention building with their own tools, such as the Kimi or GLM CLIs, but all of them mention building with Claude. There were also surprisingly few mentions of Codex, which is definitely surging in popularity in the Bay Area.Chinese companies have a technology ownership mentality. The Chinese culture is combining with a roaring economic engine to create unpredictable outcomes. I’m left with a lasting feeling that the numerous AI models reflect a practical, current equilibrium of the many technology businesses here. There’s no master plan. The industry is defined by a respect for ByteDance and Alibaba, the incumbents expected to win large portions of all markets with their substantial resources. DeepSeek is the respected technical leader, but far from a market leader. They set the direction, but aren’t set up to win economically.
This leaves companies like Meituan or Ant Group, where people in the West can be surprised they’re building these models. In reality, they see LLMs obviously as being central to future technology products, so they need a strong base. When they fine-tune the strong, general purpose model it hardens their stack from getting the open community to provide feedback on it, and they can keep internal, fine-tuned versions of the model for their products. The “open-first” mentality in the industry is largely defined by practicality — it helps make their models get strong feedback, it gives back to the open-source community, and empowers their mission.Government aid is real, but unclear how big. It’s often asserted that the Chinese government is actively helping with the open LLM race. This is a government that’s decentralized across many levels, each of which doesn’t have a clear playbook for what exactly they do. Neighborhoods in Beijing compete for tech companies to house their offices there. The “help” offered to these companies almost certainly involved removing bureaucratic red tape like permits, but how far does it go? Can levels of the government help attract talent? Can they help smuggle chips? Across the visit, there were many mentions of government interest or help, but far too little to report the details as assertive or have a confident worldview of how government can bend the trajectory of AI in China.
There were certainly no hints of the top levels of the Chinese government influencing any technical decisions in the models.The data industry is far less developed. Having heard so much about the likes of Anthropic or OpenAI spending $10M+ for single environments, with cumulative spend on the order of hundreds of millions per year to push the frontier of RL, we were eager to know if Chinese labs are either buying the same environments from companies in the U.S. or supported by a mirrored domestic ecosystem. The answer was not quite complete that there’s no data industry, but rather that their experience was that the data industry was relatively poor quality and it is often better to build the environments or data in-house. Researchers themselves spend meaningful time making the RL training environments, and some of the bigger companies like ByteDance and Alibaba can have in-house data labelling teams to support this. This all mirrors the build-not-buy mentality from the previous bullet.
Desperation for more Nvidia chips. Nvidia compute is the gold-standard for training and everyone is limited in progress by not having more of it. If supply was there, it is obvious that they would buy it. Other accelerators, including but not limited to Huawei, were spoken positively of for inference. Countless labs have access to Huawei chips.
These points paint a very different picture of an AI ecosystem, where quickly mapping how Western labs operate to their Chinese counterparts will often result in a category error. The crucial question is if these different ecosystems will produce meaningfully different types of models, or if the Chinese models will always be explained by being similar to the U.S. frontier models of 3-9 months ago.
Conclusion: The global equilibrium
I knew I knew so little about China going into the trip and came out with the feeling of just starting to learn. China isn’t a place that can be expressed by rules or recipes, but one with very different dynamics and chemistry. The culture is so old, so deep, and still completely intertwined with how domestic technology is built. I have much more learning ahead.
So much of the current power structures in the US use their current worldviews of China as crucial mental devices for decision making. Having talked, in person, either formally or informally to pretty much every leading AI lab in China, there are a lot of qualities and instincts in China that’ll be very hard to model with Western decision making. Even after asking directly about why these labs release their top models openly, the intersection between ownership mentality and genuine ecosystem support is hard for me to connect the dots on.
The labs here are practical and not necessarily absolutists around open-source, where every model they build would be released openly, but there’s a deep intentionality in supporting developers, the ecosystem, and using it as a way to learn more about their models.
Almost every major Chinese technology company is building their own general purpose LLMs, as we see with the likes of Meituan (delivery service) and Xiaomi (broad consumer technology company) releasing open weight models. The equivalent companies in the U.S. would just buy services. These companies aren’t building LLMs out of a race to be relevant with the hot new thing, but a deep fundamental yearning to control their own stack and develop the most important technologies of the day. When I look up from my laptop and always see bunches of cranes on the horizon, it obviously fits in the with the broader culture and energy around building in China.
The humanity, charm, and genuine warmth of Chinese researchers is extremely humanizing. At a personal level, the cut-throat geopolitical conversation we’re used to in the U.S. hasn’t permeated them at all. The world can use more of this simple positivity. As a citizen of the AI community, I currently worry more about the fissures appearing within members and groups around labels of nationality.
I’d be lying if I said I didn’t want US labs to be clear leaders in every part of the AI stack — especially with open models where I spend my time — I’m American, and that’s an honest preference. With this, I want the open ecosystem itself to thrive globally, as this can create safer, more accessible, and more useful AI for the world, and right now the question is whether American labs will take the steps to own that leadership position.
As of finishing this piece, more rumors are swirling of executive orders influencing open models, which can further complicate this synergy between American leadership and the global ecosystem — it doesn’t fill me with confidence.
Thank you to all the wonderful people I got to talk to at Moonshot, Zhipu, Meituan, Xiaomi, Qwen, Ant Ling, 01.ai, and others. Everyone has been so welcoming and gracious with their time. I’ll keep sharing my thoughts on China as they crystallize, across culture generally and AI specifically. It is obvious that this knowledge will be directly relevant to the story unfolding at the frontier of AI development.