超过十亿人定期使用 AI 聊天机器人。ChatGPT 每周用户超过 7 亿。Gemini 及其他领先 AI 产品又增加了数亿用户。在我的文章中,我经常聚焦于 AI 取得的进步(例如,过去几周内,OpenAI 和 Google 的 AI 聊天机器人均在国际数学奥林匹克竞赛中获得了金牌),但这掩盖了一个正在酝酿的更广泛转变:我们正在进入一个大众智能时代,强大的 AI 正变得像谷歌搜索一样触手可及。
直到最近,这些系统的免费用户(占绝大多数)只能使用较旧、较小的 AI 模型,这些模型经常出错,且对复杂工作的用途有限。最好的模型,例如能够解决极难问题且模型幻觉少得多的推理模型,需要每月支付 20 到 200 美元不等。即便如此,你还需要知道该选择哪个模型以及如何正确编写提示词。但经济性和交互界面正在迅速变化,这将对我们的工作、学习和思考方式产生相当重大的影响。
强大的 AI 正变得更便宜、更易获取
对大多数用户而言,获取强大 AI 存在两大障碍。首先是困惑。很少有人知道如何选择 AI 模型。更少人知道,在 ChatGPT 的菜单中选择 o3 能让他们使用出色的推理 AI 模型,而选择 4o(看起来数字更大)则只能得到能力远逊的模型。据 OpenAI 称,只有不到 7% 的付费客户会定期选择 o3,这意味着即使是重度用户也错过了推理模型所能做到的事情。
另一个因素是成本。由于最好的模型价格昂贵,免费用户通常无法使用它们,或者只能获得非常有限的使用权限。Google 在向免费用户提供其最佳模型方面走在前列,但 OpenAI 表示,在 GPT-5 发布之前,其几乎没有任何免费客户能定期使用推理模型。
GPT-5 本应同时解决这两个问题,这也是其首次亮相如此混乱且令人困惑的部分原因。GPT-5 实际上是两个东西。它既是一个包含多种截然不同模型的系列的总称,从较弱的 GPT-5 Nano 到强大的 GPT-5 Pro;同时,它也是用于选择使用哪个模型以及该为你的问题分配多少算力的工具的名称。当你向“GPT-5”输入内容时,你实际上是在与一个路由器对话,它应该自动判断你的问题是由一个更小、更快的模型解决,还是需要交给更强大的推理模型处理。

你可以看出这原本是为了让更多用户能使用强大的 AI:如果你只是想聊天,GPT-5 应该使用其较弱的专用聊天模型;如果你试图解决一个数学问题,GPT-5 应该将你引导至其速度较慢、成本更高的 GPT-5 思考模型。这将节省成本,并让更多人能够使用最优秀的 AI。但这次发布存在一些问题。这种做法没有得到很好的解释,而且路由器一开始工作得并不理想。结果是,使用 GPT-5 的一个人得到了非常聪明的回答,而另一个人却得到了糟糕的回答。尽管存在这些问题,OpenAI 报告称早期取得了成功。在发布后的几天内,使用过推理模型的付费用户比例从 7% 上升到了 24%,而使用最强大模型的免费用户比例则从几乎为零上升到了 7%。
这种变化的部分原因在于,更智能的模型运行效率正在大幅提升。这张图展示了这一趋势的发展速度,y 轴表示 AI 的能力,x 轴表示对数递减的成本。当 GPT-4 推出时,处理一百万个 token(一个 token 大约相当于一个单词)的成本约为 50 美元,而现在,使用比原始 GPT-4 强大得多的 GPT-5 nano 模型,每百万 token 的成本约为 14 美分。

这种效率提升不仅体现在经济上,也体现在环境方面。谷歌报告称,仅在过去一年里,每次提示词的能量效率就提高了33倍。根据独立测试和官方公告,2025年,现代大语言模型处理一个标准提示词所消耗的边际能量已相对确定。大约为0.0003千瓦时,相当于流媒体播放Netflix 8-10秒的能耗,或2008年一次谷歌搜索的能耗(有趣的是,图像生成似乎与文本提示词消耗的能量相近)。这些模型每次提示词消耗多少水则不太明确,根据用水定义的不同,范围从几滴水到五分之一小酒杯(0.25毫升到5毫升以上)不等(这里是低耗水量的论点,这里是高耗水量的论点)。
这些改进意味着,即使人工智能变得更强大,它也能为更多人提供服务。服务每个新增用户的边际成本已经大幅下降,这使得更多商业模式(如广告支持)成为可能。免费用户现在可以运行那些仅仅两年前还需要花费数美元的提示词。这就是十亿人突然能够使用强大AI的方式:并非通过某种宏大的民主化倡议,而是因为经济条件最终使之成为可能。
强大的人工智能正变得易于使用
获得强大 AI 的访问权限还不够,人们需要真正用它来完成工作。过去,用好 AI 是一个相当有挑战性的过程,需要运用链式思维等技巧来精心设计提示词,同时学习各种技巧和窍门,才能充分发挥 AI 的能力。然而,在最近的一系列实验中,我们发现这些技巧已经不再那么有用了。强大的 AI 模型越来越擅长执行你的指令,甚至能揣摩你的意图,并超越你的要求(而且,平均而言,威胁它们或对它们友善似乎也没什么帮助)。
而且,变得更便宜、更易用的不仅仅是文本模型。谷歌发布了一款代号为“nano banana”、官方名称则平淡得多的 Gemini 2.5 Flash Image Generator 的新图像模型。这款模型不仅表现出色(尽管更擅长编辑图像而非创建新图像),而且价格足够低廉,免费用户也能使用。与之前几代 AI 图像生成器不同,它能够很好地遵循自然语言指令。
为了展示其强大功能和易用性,我上传了一张标志性的(且无版权问题的)阿波罗 11 号宇航员照片和一张闪亮燕尾服的随机图片,并给出了最简单的提示词:“给左边的尼尔·阿姆斯特朗穿上这件燕尾服”。
几秒钟后,它给出了以下结果:
虽然专业人士一眼就能看出一些问题,但燕尾服逼真的褶皱以及它如何融入场景(翻领上的 NASA 徽章是个不错的细节)仍然令人印象深刻。目前该过程仍存在大量随机性,使得 AI 图像编辑不适合许多专业应用,但对大多数人来说,这代表着巨大的飞跃——不仅在于他们能做什么,更在于做起来有多容易。
我们还可以更进一步:“现在展示一张照片,尼尔·阿姆斯特朗和巴兹·奥尔德林穿着同样的服装,坐在现代飞机的座位上,尼尔看起来很放松,向后靠着,正在吹小号,巴兹看起来很紧张,手里拿着一个汉堡,中间的座位上坐着一只逼真的水獭,它坐在座位上,正在使用笔记本电脑。”
这包含多重含义:AI 产出了相当令人印象深刻的成果(看看那些表情,以及它如何保留了巴兹的戒指和尼尔的领针)。这是 AI 对一个著名历史时刻的扭曲再现。同时,这也是一种潜在的警示,提醒我们当这类技术被广泛使用时,事情会变得多么怪异。
大众智能的怪异之处
当强大的 AI 掌握在十亿人手中时,很多事情会同时发生。很多事情已经在同时发生了。
有些人正与 AI 模型建立深厚的关系,而另一些人则因此摆脱了孤独。AI 模型可能正在导致一些人精神崩溃和危险行为,同时又被用来诊断其他人的疾病。它被用来撰写讣告、创作经文、在作业上作弊、创办新企业,以及成千上万种其他意想不到的用途。随着 AI 系统变得更加强大,这些用途,以及随之而来的问题和益处,很可能只会成倍增加。
尽管谷歌的 AI 图像生成器设有防止滥用的护栏,以及用于识别 AI 图像的无形水印,但我预计,在未来几个月内,限制少得多的 AI 图像生成器很可能在质量上接近纳米香蕉的水平。
AI 公司(无论你是否相信它们对安全的承诺)似乎和我们其他人一样,无法完全消化这一切。当十亿人都能使用先进的 AI 时,我们就进入了可以称之为大众智能的时代。我们现有的每一个机构——学校、医院、法院、公司、政府——都是为一个智能稀缺且昂贵的世界而建立的。现在,每一个行业、每一个机构、每一个社区都必须弄清楚如何在大众智能时代蓬勃发展。我们如何驾驭十亿人使用 AI 的局面,同时管理随之而来的混乱?当任何人都可以伪造任何东西时,我们如何重建信任?我们如何在普及知识获取渠道的同时,保留人类专业知识中宝贵的东西?
所以,我们走到了这一步。强大的 AI 已经廉价到可以免费赠送,简单到无需说明书,并且能力足以在多种智力任务上超越人类。一波机遇与问题即将涌现在全球的教室、法庭和董事会会议室。大众智能时代,就是当十亿人获得前所未有的工具集,并观察他们如何运用时,所发生的一切。我们即将亲身体验那会是怎样一番景象。
这是回答一个标准提示词所需的能量。它并未计入训练 AI 模型所需的能量——那是一次性的、极其耗能的过程。我们不知道构建一个现代模型究竟消耗了多少能量,但据估算,训练 GPT-4 大约需要 50 万千瓦时,相当于一架波音 737 飞行约 18 小时。
More than a billion people use AI chatbots regularly. ChatGPT has over 700 million weekly users. Gemini and other leading AIs add hundreds of millions more. In my posts, I often focus on the advances that AI is making (for example, in the past few weeks, both OpenAI and Google AIs chatbots got gold medals in the International Math Olympiad), but that obscures a broader shift that's been building: we're entering an era of Mass Intelligence, where powerful AI is becoming as accessible as a Google search.
Until recently, free users of these systems (the overwhelming majority) had access only to older, smaller AI models that frequently made mistakes and had limited use for complex work. The best models, like Reasoners that can solve very hard problems and hallucinate much less often, required paying somewhere between $20 and $200 a month. And even then, you needed to know which model to pick and how to prompt it properly. But the economics and interfaces are changing rapidly, with fairly large consequences for how all of us work, learn, and think.
Powerful AI is Getting Cheaper and Easier to Access
There have been two barriers to accessing powerful AI for most users. The first was confusion. Few people knew to select an AI model. Even fewer knew that picking o3 from a menu in ChatGPT would get them access to an excellent Reasoner AI model, while picking 4o (which seems like a higher number) would give them something far less capable. According to OpenAI, less than 7% of paying customers selected o3 on a regular basis, meaning even power users were missing out on what Reasoners could do.
Another factor was cost. Because the best models are expensive, free users were often not given access to them, or else given very limited access. Google led the way in giving some free access to its best models, but OpenAI stated that almost none of its free customers had regular access to reasoning models prior to the launch of GPT-5.
GPT-5 was supposed to solve both of these problems, which is partially why its debut was so messy and confusing. GPT-5 is actually two things. It was the overall name for a family of quite different models, from the weaker GPT-5 Nano to the powerful GPT-5 Pro. It was also the name given to the tool that picked which model to use and how much computing power the AI should use to solve your problem. When you are writing to “GPT-5” you are actually talking to a router that is supposed to automatically decide whether your problem can be solved by a smaller, faster model or needs to go to a more powerful Reasoner.

You could see how this was supposed to expand access to powerful AI to more users: if you just wanted to chat, GPT-5 was supposed to use its weaker specialized chat models; if you were trying to solve a math problem, GPT-5 was supposed to send you to its slower, more expensive GPT-5 Thinking model. This would save money and give more people access to the best AIs. But the rollout had issues. This practice wasn’t well explained and the router did not work well at first. The result is that one person using GPT-5 got a very smart answer while another got a bad one. Despite these issues, OpenAI reported early success. Within a few days of launch, the percentage of paying customers who had used a Reasoner went from 7% to 24% and the number of free customers using the most powerful models went from almost zero to 7%.
Part of this change is driven by the fact that smarter models are getting dramatically more efficient to run. This graph shows how fast this trend has played out, mapping the capability of AI on the y-axis and the logarithmically decreasing costs on the x-axis. When GPT-4 came out it was around $50 to work with a million tokens (a token is roughly a word), now it costs around 14 cents per million tokens to use GPT-5 nano, a much more capable model than the original GPT-4.

This efficiency gain isn't just financial, it's also environmental. Google has reported that energy efficiency per prompt has improved by 33x in the last year alone. The marginal energy used by a standard prompt from a modern LLM in 2025 is relatively established at this point, from both independent tests and official announcements. It is roughly 0.0003 kWh, the same energy use as 8-10 seconds of streaming Netflix or the equivalent of a Google search in 2008 (interestingly, image creation seems to use a similar amount of energy as a text prompt)1. How much water these models use per prompt is less clear but ranges from a few drops to a fifth of a shot glass (.25mL to 5mL+), depending on the definitions of water use (here is the low water argument and the high water argument).
These improvements mean that even as AI gets more powerful, it's also becoming viable to give to more people. The marginal cost of serving each additional user has collapsed, which means more business models, like ad support, become possible. Free users can now run prompts that would have cost dollars just two years ago. This is how a billion people suddenly get access to powerful AIs: not through some grand democratization initiative, but because the economics finally make it possible.
Powerful AI is Getting Easy to Use
Getting access to a powerful AI is not enough, people need to actually use it to get things done. Using AI well used to be a pretty challenging process which involved crafting a prompt using techniques like chain-of-thought along with learning tips and tricks to get the most out of your AI. In a recent series of experiments, however, we have discovered that these techniques don’t really help anymore. Powerful AI models are just getting better at doing what you ask them to or even figuring out what you want and going beyond what you ask (and no, threatening them or being nice to them does not seem to help on average).
And it isn’t just text models that are becoming cheaper and easier to use. Google released a new image model with the code name “nano banana” and the much more boring official name Gemini 2.5 Flash Image Generator. In addition to being excellent (though better at editing images than creating new ones), it is also cheap enough that free users can access it. And, unlike previous generations of AI image generators, it follows instructions in plain language very well.
As an example of both its power and ease of use, I uploaded an iconic (and copyright free) image of the Apollo 11 astronauts and a random picture of a sparkly tuxedo and gave it the simplest prompts: “dress Neil Armstrong on the left in this tuxedo”
Here is what it gave me a few seconds later:
There are issues that someone with an expert eye would spot, but it is still impressive to see the realistic folds of the tuxedo and how it is blended into the scene (the NASA pin on the lapel was a nice touch). There is still a lot of randomness in the process that makes AI image editing unsuitable for many professional applications, but for most people, this represents a huge leap in not just what they can do, but how easy it is to do it.
And we can go further: “now show a photograph where neil armstrong and buzz aldrin, in the same outfits, are sitting in their seats in a modern airplane, neil looks relaxed and is leaning back, playing a trumpet, buzz seems nervous and is holding a hamburger, in the middle seat is a realistic otter sitting in a seat and using a laptop.”
This is many things: A pretty impressive output from the AI (look at the expressions, and how it preserved Buzz’s ring and Neil’s lapel pin). A distortion of a famous moment in history made possible by AI. And a potential warning about how weird things are going to get when these sorts of technologies are used widely.
The Weirdness of Mass Intelligence
When powerful AI is in the hands of a billion people, a lot of things are going to happen at once. A lot of things are already happening at once.
Some people have intense relationships with AI models while other people are being saved from loneliness. AI models may be causing mental breakdowns and dangerous behavior for some while being used to diagnose the diseases of others. It is being used to write obituaries and create scriptures and cheat on homework and launch new ventures and thousands of other unexpected uses. These uses, and both the problems and benefits, are likely to only multiply as AI systems get more powerful.
And while Google's AI image generator has guardrails to limit misuse, as well as invisible watermarks to identify AI images, I expect much less restrictive AI image generators will likely get close to nano banana in quality in the coming months.
The AI companies (whether you believe their commitments to safety or not) seem to be as unable to absorb all of this as the rest of us are. When a billion people have access to advanced AI, we've entered what we might call the era of Mass Intelligence. Every institution we have — schools, hospitals, courts, companies, governments — was built for a world where intelligence was scarce and expensive. Now every profession, every institution, every community has to figure out how to thrive with Mass Intelligence. How do we harness a billion people using AI while managing the chaos that comes with it? How do we rebuild trust when anyone can fabricate anything? How do we preserve what's valuable about human expertise while democratizing access to knowledge?
So here we are. Powerful AI is cheap enough to give away, easy enough that you don't need a manual, and capable enough to outperform humans at a range of intellectual tasks. A flood of opportunities and problems are about to show up in classrooms, courtrooms, and boardrooms around the world. The Mass Intelligence era is what happens when you give a billion people access to an unprecedented set of tools and see what they do with it. We are about to find out what that is like.
This is the energy required to answer a standard prompt. It does not take into account the energy needed to train AI models, which is a one-time process that is very energy intensive. We do not know how much energy is used to create a modern model, but it was estimated that training GPT-4 took a little above 500,000 kWh, about 18 hours of a Boeing 737 in flight.