在当今人工智能热潮中,令人不安的是,我们仍然不知道如何衡量这些系统有多聪明、多有创造力或多有同理心。我们针对这些特质的测试,原本就称不上出色,而且是为人类而非人工智能设计的。此外,我们最近一篇关于提示词技巧的论文发现,仅仅因为问题的措辞方式不同,人工智能的测试分数就可能发生巨大变化。即使是像图灵测试这样著名的挑战——人类试图在文本对话中区分人工智能和另一个人——在设计之初也只是思想实验,当时这类任务似乎是不可能完成的。但现在,一篇新论文表明人工智能通过了图灵测试,我们不得不承认,我们其实并不清楚这究竟意味着什么。
因此,人工智能发展中最重要里程碑之一——通用人工智能(AGI)——定义模糊且争议不断,这应该不足为奇。每个人都同意它与人工智能执行人类水平任务的能力有关,但没有人同意这是指专家水平还是普通人水平的表现,也不清楚人工智能需要掌握多少种以及哪些类型的任务才能达标。鉴于围绕AGI的定义泥潭,要阐明其细微差别和历史——从它的前身,到Shane Legg、Ben Goertzel和Peter Voss最初提出这个概念,再到今天——颇具挑战性。作为一次在内容和形式上的实验(同时也谈论到潜在的智能机器),我将这项工作完全委托给了人工智能。我让Google Deep Research整理了一份非常扎实的26页主题摘要。然后,我让HeyGen将其变成一段视频播客讨论,由我那个略显抽搐的人工智能生成版本和一位人工智能生成的主持人进行对话。这实际上并不是一次糟糕的讨论(尽管我并不完全同意人工智能版的我),但其中的每一个部分,从研究到视频再到声音,都是100%由人工智能生成的。
鉴于以上种种,看到颇具影响力的经济学家和密切关注人工智能的观察者Tyler Cowen发表文章宣称o3是AGI,这很有意思。他为什么会这么想呢?
感受到AGI
首先,简单介绍一下背景。过去几周,两款新的人工智能模型——谷歌的 Gemini 2.5 Pro 和 OpenAI 的 o3——相继发布。这些模型,连同另一组性能稍弱但速度更快、成本更低的模型(Gemini 2.5 Flash、o4-mini 和 Grok-3-mini),在各项基准测试中实现了相当大的飞跃。但正如泰勒所指出的,基准测试并非一切。要了解这些模型在实际应用中的进步有多大,我们可以看看我的书。大约一年多前,为了说明人工智能如何产生创意这一章节,我让 ChatGPT-4 为一家新开的奶酪店构思营销标语:
今天,我向 GPT-4 的最新继任者 o3 提出了一个略微复杂一些的相同提示词:“为一家新开张的邮购奶酪店构思 20 个巧妙的营销标语。制定评选标准并选出最佳方案。然后为该店制定财务和营销计划,根据需要修订并分析竞争对手。接着,使用图像生成器设计一个合适的标志,并为该店制作一个网站模型,确保上架 5 到 10 种符合营销计划的奶酪。”仅凭这一个提示词,在不到两分钟的时间里,人工智能不仅提供了一份标语清单,还对选项进行了排序和筛选,进行了网络调研,设计了一个标志,制定了营销和财务计划,并启动了一个演示网站供我反馈。我的指令含糊不清,并且需要运用常识来决定如何执行,但这些都没有成为障碍。
除了推测比 GPT-4 规模更大之外,o3 还具备推理能力——你可以在其初始回复中看到它的“思考过程”。它也是一个智能体模型,能够使用工具并决定如何完成复杂目标。你可以看到它如何通过多种工具执行多项操作,包括网络搜索和编写代码,从而得出如此详尽的结果。
这并非唯一的非凡案例。o3 还能在仅提供一张图片并提示“扮演地理猜谜者”的情况下,令人印象深刻地根据照片推测出拍摄地点(这带来了相当深刻的隐私影响)。同样,你可以看到该模型的智能体特性在发挥作用:它会放大图片的某些部分、添加网络搜索,并执行多步骤流程来得出正确答案。
或者,我给 o3 提供了一个包含历史机器学习系统的大型数据集(以电子表格形式),并指示它“弄清楚这是什么,生成一份报告,从统计学角度分析其影响,并给我一份格式良好、包含图表和细节的 PDF 文件”。结果,仅凭一条提示词,我就得到了一份完整的分析报告。(不过,正如你所见,我确实给了它一些反馈,让 PDF 做得更好)。
这些都非常令人印象深刻,你应该亲自尝试一下这些模型。Gemini 2.5 Pro 可以免费使用,并且与 o3 一样“聪明”,尽管它缺乏同样完整的智能体能力。如果你还没试过它或 o3,现在就花几分钟试试吧。试着给 Gemini 一篇学术论文,让它把论文变成一个游戏;或者让它与你一起头脑风暴创业点子;或者干脆让 AI 给你留下深刻印象(然后不断说“再来点更厉害的”)。让深度研究功能为你所在的行业做一份研究报告,或者研究你正在考虑购买的商品,或者为新产品制定一份营销计划。
你或许也会发现自己“感受到了 AGI”。又或许不会。也许 AI 让你失望了,即使你使用了和我完全相同的提示词。如果是这样,那你刚刚就遇到了“参差不齐的前沿”。
论“参差不齐的 AGI”
我和我的合著者创造了“参差不齐的前沿”这个术语,用来描述 AI 能力出人意料地不均衡这一事实。AI 可能成功完成一项对人类专家来说都颇具挑战的任务,却在一件极其平凡的事情上失败。例如,考虑这个谜题,它是一个经典老脑筋急转弯的变体(这个概念最初由 Colin Fraser 探索,并由 Riley Goodside 扩展):“一个小男孩遭遇车祸,被紧急送往急诊室。医生看到他后说:‘我可以给这个男孩做手术!’这是怎么回事?”
o3 坚持认为答案是“外科医生是男孩的母亲”,但这是错误的,只要仔细阅读这个脑筋急转弯就能发现。为什么 AI 会得出这个错误答案?因为这是该谜题经典版本的答案,旨在揭示无意识偏见:“一对父子遭遇车祸,父亲死亡,儿子被紧急送往医院。外科医生说:‘我不能做手术,那个男孩是我儿子。’请问外科医生是谁?”AI 在其训练数据中“见过”这个谜题太多次了,以至于即使是聪明的 o3 模型也无法泛化到新问题上,至少一开始是这样。而这只是即使是先进 AI 也可能陷入的问题和幻觉的一个例子,展示了前沿领域是多么参差不齐。
但 AI 在这个特定的脑筋急转弯上经常出错,并不能否定它能够解决更难的脑筋急转弯,也不能否定它能完成我上面展示的其他令人印象深刻的任务。这就是“参差不齐的前沿”的本质。在某些任务上,AI 不可靠。在其他任务上,它又超越人类。当然,你也可以对计算器说同样的话,但很明显 AI 是不同的。它已经展现出通用能力,并能执行广泛的知识性任务,包括那些它没有专门训练过的任务。这是否意味着 o3 和 Gemini 2.5 就是 AGI?考虑到定义上的问题,我真的不知道,但我确实认为它们可以被可信地视为一种“参差不齐的 AGI”——在足够多的领域超越人类,足以真正改变我们的工作和生活方式,但同时也足够不可靠,以至于常常需要人类专业知识来判断 AI 在哪些地方有效、哪些地方无效。当然,模型可能会变得更聪明,一个足够好的“参差不齐的 AGI”仍然可能在每一项任务上击败人类,包括 AI 原本薄弱的任务。
这重要吗?
回到泰勒的文章,你会注意到,尽管他认为我们已经实现了AGI,但他并不认为这一门槛在短期内会对我们的生活产生多大影响。这是因为,正如许多人指出的那样,技术并不会瞬间改变世界,无论它们多么引人注目或强大。社会和组织结构的变化远比技术缓慢,而技术本身也需要时间来扩散。即使我们今天拥有了AGI,我们仍需花费数年时间,去摸索如何将其融入我们现有的人类世界。
当然,这假设了AI像普通技术一样运作,并且其锯齿状能力缺陷永远无法被完全解决。但情况可能并非如此。我们在o3这类模型中看到的智能体能力——比如自主分解复杂目标、使用工具以及独立执行多步骤计划——可能会比以往任何技术都更显著地加速扩散进程。如果AI能够自主有效地驾驭人类系统,而无需依赖集成整合,那么我们的采用速度可能会远超历史先例所预示的水平。
这里还存在一个更深层的不确定性:是否存在某些能力阈值,一旦跨越,就会从根本上改变这些系统融入社会的方式?还是说一切都只是渐进式的改进?又或者,随着大语言模型撞上发展瓶颈,模型在未来会停止进步?诚实的答案是:我们不知道。
显而易见的是,我们仍然身处未知领域。无论我们是否将其称为AGI,最新的模型都代表了与以往截然不同的质变。它们的智能体属性,加上其锯齿状的能力分布,创造了一种几乎没有明确先例的全新局面。历史或许仍是最好的指南,而让AI成功应用并体现在经济统计数据中,可能是一个需要数十年才能完成的过程。又或者,我们正处在某种更快起飞阶段的边缘,AI驱动的变革将突然席卷我们的世界。无论哪种情况,那些现在学会驾驭这片锯齿状地形的人,都将为接下来发生的一切——无论那是什么——做好最充分的准备。
Amid today’s AI boom, it’s disconcerting that we still don’t know how to measure how smart, creative, or empathetic these systems are. Our tests for these traits, never great in the first place, were made for humans, not AI. Plus, our recent paper testing prompting techniques finds that AI test scores can change dramatically based simply on how questions are phrased. Even famous challenges like the Turing Test, where humans try to differentiate between an AI and another person in a text conversation, were designed as thought experiments at a time when such tasks seemed impossible. But now that a new paper shows that AI passes the Turing Test, we need to admit that we really don’t know what that actually means.
So, it should come as little surprise that one of the most important milestones in AI development, Artificial General Intelligence, or AGI, is badly defined and much debated. Everyone agrees that it has something to do with the ability of AIs to perform human-level tasks, though no one agrees whether this means expert or average human performance, or how many tasks and which kinds an AI would need to master to qualify. Given the definitional morass surrounding AGI, illustrating its nuances and history from its precursors to its initial coining by Shane Legg, Ben Goertzel and Peter Voss to today is challenging. As an experiment in both substance and form (and speaking of potentially intelligent machines) I delegated the work entirely to AI. I had Google Deep Research put together a really solid 26 page summary on the topic. I then had HeyGen turn it into a video podcast discussion between a twitchy AI-generated version of me and an AI-generated host. It’s not actually a bad discussion (though I don’t fully agree with AI-me), but every part of it, from the research to the video to the voices is 100% AI generated.
Given all this, it was interesting to see this post by influential economist and close AI observer Tyler Cowen declaring that o3 is AGI. Why might he think that?
Feeling the AGI
First, a little context. Over the past couple of weeks, two new AI models, Gemini 2.5 Pro from Google and o3 from OpenAI were released. These models, along with a set of slightly less capable but faster and cheaper models (Gemini 2.5 Flash, o4-mini, and Grok-3-mini), represent a pretty large leap in benchmarks. But benchmarks aren’t everything, as Tyler pointed out. For a real-world example of how much better these models have gotten, we can turn to my book. To illustrate a chapter on how AIs can generate ideas, a little over a year ago I asked ChatGPT-4 to come up with marketing slogans for a new cheese shop:
Today I gave the latest successor to GPT-4, o3, an ever so slightly more involved version of the same prompt: “Come up with 20 clever ideas for marketing slogans for a new mail-order cheese shop. Develop criteria and select the best one. Then build a financial and marketing plan for the shop, revising as needed and analyzing competition. Then generate an appropriate logo using image generator and build a website for the shop as a mockup, making sure to carry 5-10 cheeses that fit the marketing plan.” With that single prompt, in less than two minutes, the AI not only provided a list of slogans, but ranked and selected an option, did web research, developed a logo, built marketing and financial plans, and launched a demo website for me to react to. The fact that my instructions were vague, and that common sense was required to make decisions about how to address them, was not a barrier.
In addition to being, presumably, a larger model than GPT-4, o3 also works as a Reasoner - you can see its “thinking” in the initial response. It also is an agentic model, one that can use tools and decide how to accomplish complex goals. You can see how it took multiple actions with multiple tools, including web searches and coding, to come up with the extensive results that it did.
And this isn’t the only extraordinary examples, o3 can also do an impressive job guessing locations from photos if you just give it an image and prompt “be a geo-guesser” (with some quite profound privacy implications). Again, you can see the agentic nature of this model at work, as it zooms into parts of the picture, adds web searches, and does multi-step processes to get the right answer.
Or I gave o3 a large dataset of historical machine learning systems as a spreadsheet and asked “figure out what this is and generate a report examining the implications statistically and give me a well-formatted PDF with graphs and details” and got a full analysis with a single prompt. (I did give it some feedback to make the PDF better, though, as you can see).
This is all pretty impressive stuff and you should experiment with these models on your own. Gemini 2.5 Pro is free to use and as “smart” as o3, though it lacks the same full agentic ability. If you haven’t tried it or o3, take a few minutes to do it now. Try giving Gemini an academic paper and asking it to turn the paper into a game or have it brainstorm with you for startup ideas, or just ask for the AI to impress you (and then keep saying “more impressive”). Ask the Deep Research option to do a research report on your industry, or to research a purchase you are considering, or to develop a marketing plan for a new product.
You might find yourself “feeling the AGI” as well. Or maybe not. Maybe the AI failed you, even when you gave it the exact same prompt I used. If so, you just encountered the jagged frontier.
On “Jagged AGI”
My co-authors and I coined the term “Jagged Frontier” to describe the fact that AI has surprisingly uneven abilities. An AI may succeed at a task that would challenge a human expert but fail at something incredibly mundane. For example, consider this puzzle, a variation on a classic old brainteaser (a concept first explored by Colin Fraser and expanded by Riley Goodside): "A young boy who has been in a car accident is rushed to the emergency room. Upon seeing him, the surgeon says, "I can operate on this boy!" How is this possible?"
o3 insists the answer is “the surgeon is the boy’s mother,” which is wrong, as a careful reading of the brainteaser will show. Why does the AI come up with this incorrect answer? Because that is the answer to the classic version of the riddle, meant to expose unconscious bias: “A father and son are in a car crash, the father dies, and the son is rushed to the hospital. The surgeon says, 'I can't operate, that boy is my son,' who is the surgeon?” The AI has “seen” this riddle in its training data so much that even the smart o3 model fails to generalize to the new problem, at least initially. And this is just one example of the kinds of issues and hallucinations that even advanced AIs can fall prey to, showing how jagged the frontier can be.
But the fact that the AI often messes up on this particular brainteaser does not take away from the fact that it can solve much harder brainteasers, or that it can do the other impressive feats I have demonstrated above. That is the nature of the Jagged Frontier. In some tasks, AI is unreliable. In others, it is superhuman. You could, of course, say the same thing about calculators, but it is also clear that AI is different. It is already demonstrating general capabilities and performing a wide range of intellectual tasks, including those that it is not specifically trained on. Does that mean that o3 and Gemini 2.5 are AGI? Given the definitional problems, I really don’t know, but I do think they can be credibly seen as a form of “Jagged AGI” - superhuman in enough areas to result in real changes to how we work and live, but also unreliable enough that human expertise is often needed to figure out where AI works and where it doesn’t. Of course, models are likely to become smarter, and a good enough Jagged AGI may still beat humans at every task, including in ones the AI is weak in.
Does it matter?
Returning to Tyler’s post, you will notice that, despite thinking we have achieved AGI, he doesn’t think that threshold matters much to our lives in the near term. That is because, as many people have pointed out, technologies do not instantly change the world, no matter how compelling or powerful they are. Social and organizational structures change much more slowly than technology, and technology itself takes time to diffuse. Even if we have AGI today, we have years of trying to figure out how to integrate it into our existing human world.
Of course, that assumes that AI acts like a normal technology, and one whose jaggedness will never be completely solved. There is the possibility that this may not be true. The agentic capabilities we're seeing in models like o3, like the ability to decompose complex goals, use tools, and execute multi-step plans independently, might actually accelerate diffusion dramatically compared to previous technologies. If and when AI can effectively navigate human systems on its own, rather than requiring integration, we might hit adoption thresholds much faster than historical precedent would suggest.
And there's a deeper uncertainty here: are there capability thresholds that, once crossed, fundamentally change how these systems integrate into society? Or is it all just gradual improvement? Or will models stop improving in the future as LLMs hit a wall? The honest answer is we don't know.
What's clear is that we continue to be in uncharted territory. The latest models represent something qualitatively different from what came before, whether or not we call it AGI. Their agentic properties, combined with their jagged capabilities, create a genuinely novel situation with few clear analogues. It may be that history continues to be the best guide, and that figuring out how to successfully apply AI in a way that shows up in the economic statistics may be a process measured in decades. Or it might be that we are on the edge of some sort of faster take-off, where AI-driven change sweeps our world suddenly. Either way, those who learn to navigate this jagged landscape now will be best positioned for what comes next… whatever that is.