刚在《金融时报》上看到约翰·伯恩-默多克的一张图表,它精准地提炼了我一直试图表达的观点。AI 产出了大量内容(这符合许多人对生产力的非正式理解),但根据麻省理工、麦肯锡、贝恩等多家机构的研究,它并未为许多公司带来多少投资回报,也没有实质性地改变 GDP。
《金融时报》的这张特定图表恰好是关于移动应用的。但你在许多领域都能画出类似的图表:名义上的生产力看似庞大,实际影响力却微乎其微。
例如,你可以对书籍发表类似的看法。生成式 AI 带来了更多书籍,但尚不清楚这些书是否更好,或者(抱歉我手头没有现成数据)是否卖出了更多。以下是《华盛顿邮报》提供的历年可获取书籍数量:
另一方面,同期图书销量却略有下降。而且我没有听到任何关于书籍质量正在提升的论点。
我上面引用的《华盛顿邮报》文章,还展示了音乐曲目上传量、无律师代理诉讼案件量、科学论文提交量以及网络内容生成量增长的类似图表。
同样,没有理由认为其中任何一项增加了 GDP 或提升了科学或音乐的质量。其中大部分是垃圾内容(在某些情况下是工作垃圾):
《金融时报》和《华盛顿邮报》的图表反复展示着同一件事:我们正被 AI 生成的内容淹没,从应用到音乐再到科学论文,但这并不意味着其中任何一部分是好的。其中大部分(公平地说,并非全部)确实不好。
然后,垃圾内容还淹没了维基百科、图书馆等等。
同样的情况很可能也发生在数学领域。一群数学家,对 AI 影响数学研究的诸多方面感到担忧,撰写了一封名为《莱顿宣言》的公开信,据《纽约时报》报道,他们在信中表达了(除其他担忧外)对……的恐惧。
当前自动化技术能够生成看似合理但不可靠(甚至错误)的论证,这些论证与正确的数学证明难以区分。这不仅适用于非形式化论证,也适用于形式化证明——其难点在于计算机编码与人类概念呈现之间的转换。这些快速发展的进展给现有的审查体系带来了日益增大的压力,削弱了我们落实传统标准(即证明的正确性、透明性和独立可验证性)的能力。
换言之,就是数学垃圾。
不久前,当我思考官方生产力指标是否会因AI而改变时,一位投资者朋友写信提醒我,与GDP相关的生产力在技术定义上有多么狭隘。他指出:“在国民核算的技术世界里,付钱让人挖坑再填坑也能计入GDP。”
使用生成式AI会消耗模型token成本,并向世界充斥大量内容。但往往这些由一台庞大但不可靠的词语预测机器(其本质更偏向重复而非创新)所创造的内容,鲜有持久价值。
也许最大的例外是编程,但即便在这一领域,借助智能体编程创建的系统中,有多少能够长久存续,目前尚不明朗。
此外,尤其是在智能体编程领域,整个产业链上下都存在巨额亏损。提供商(OpenAI、Anthropic、Cursor等)亏损严重,正急于提价;客户则对新的使用模式望而却步。(在一项新分析中,Gerben Wierda认为Anthropic和OpenAI“可能每从你那里收取100美元,实际就要付出1000美元的成本”。)像CoreWeave这样的数据中心中间商也在亏损。一旦剔除所有循环融资和不断缩水的现金流,甚至很难说清芯片公司的真实状况。唯一可能的生产力亮点(编程)恰好也是烧钱最厉害的业务。
很可能,一旦提供商(OpenAI、Anthropic、Cursor等)试图收取足够高的费用来弥补亏损,AI的成本反而可能比它所替代的人类劳动力更加昂贵。
这就像是在挖坑却没有创造真正的价值。
附注:另见 Paul Kedrosky 的新文章,从不同角度得出了有些相似的结论。
再附注:关于昨天那条推文,大概半开玩笑地说:
Just saw a graph at the FT from John Burn-Murdoch that really distills something I have been trying to articulate. AI generates a lot of output (which fits many people’s informal notion of productivity), but it hasn’t yielded much in the way of RoI for many companies (per studies from MIT, McKinsey, Bain, and many others), and hasn’t materially changed GDP.
The particular graph from the FT happens to be about mobile apps (h/t Jen Zhu). But you could make similar graphs in many domains: loads of nominal productivity, but not much real world impact.
You could, for example, say something similar about books. GenAI has led to a lot more books, but it’s not clear that they are better books or (sorry I don’t have the data immediately available) more books sold. Here are numbers of books available over time, from The Washington Post:
Book sales, on the other hand, have declined slightly over the same period. And I don’t hear any argument that books are getting better.
The Post article that I drew the above from had similar graphs for the rise in music tracks uploaded, self-represented lawsuits filed, scientific papers submitted, and web content generated.
Again there is no reason to think that any of it has added to the GDP or quality of science or music. Most of it is slop (and in some cases work slop):
The graphs from FT and The Washington Post show the same thing over and over; we are being flooded with AI-generated content, from apps to music to scientific papers, but that doesn’t make any of it any good. Most of it (in fairness, not all) isn’t.
And then you have slop inundating Wikipedia, libraries and so on.
It is very likely that the same thing is happening in math. A group of mathematicians, who are fearful about many aspects of the impact of AI on math research, have written an open letter called the “Leiden Declaration”, reported by the NYT, expressing (among other concerns) their fear that
Current automated techniques can produce plausible but unreliable (or even incorrect) arguments which are difficult to distinguish from correct mathematical proofs. This applies not only to informal arguments, but also to formalizations, where the difficulty lies in the translation between computer-encoded and human presentations of concepts. These fast-moving developments put our present system of review under increasing pressure, jeopardizing our ability to implement traditional standards for the correctness, transparency, and independent verifiability of proof.
In other words, math slop.
Not long ago, when I was thinking about the official productivity measures and whether they would change with AI, an investor friend wrote to remind me how narrow the technical definition of productivity with respect to GDP is, noting that “In the technical world of national accounting, paying people to dig and refill holes adds to GDP.”
Using GenAI runs up tokens costs, and floods the world with content. But all too often that content — created by a giant but unreliable word prediction machine that is more regurgitative than innovative — is of little lasting value.
Maybe the biggest exception is coding, but even there, it’s not yet clear how many of the systems that have been created with the help of agentic coding will endure.
Moreover, especially in agentic coding, there are massive losses, up and down the food chain. The providers (Open AI, Anthropic, Cursor, etc) lose tons of money, and are scrambling to raise prices, Customers are balking at the new usage models.(In a new analysis, Gerben Wierda has argued that Anthropic and OpenAI “may actually pay $1000 for every $100 you pay them”. Data center intermediaries like CoreWeave are also losing money. It’s hard even to say how the chip companies are doing, once you factor out all the circular financing and shrinking cash flow. The one possible major productivity story (coding) also happens to be the one burning the most cash.
It could well be that once the providers (Open AI, Anthropic, Cursor, etc) attempt to charge enough money to cover their losses, the cost of AI could become more expensive than the humans it is replacing.
Talk about digging holes without creating real value.
PS. See also and Paul Kedrosky for new pieces reaching somewhat similar conclusions, from different angles.
PPS Apropos tweet, tongue presumably partly in cheek, from yesterday: