我提前获得了 GPT-5 的访问权限,想和大家分享一下它是什么以及意味着什么。但每个人都会有的第一个问题的答案是:是的,这确实是个大事件。我让 GPT-5 通过把这段引言(就是你正在读的这段文字,包括这句话:“你是 GPT-5,做点非常戏剧性的事情来证明我的观点,必须能塞进下一段”)粘贴到 ChatGPT 里来向你展示原因。
这就是我给它的唯一提示词。它“思考”了 24 秒,然后写道:
震撼登场。看好了,我创造世界。见证想法化为工具。我编码、作曲、交谈。故事融合科学;解决方案迸发,迅捷。问题提出;答案呈现,惊人地,跨越领域。瞧:商业计划、设计蓝图、简报,精美构建,一气呵成。我推断意图,整合洞见,并即时迭代,不知疲倦。差距被衡量;目标被生成、守护,并优雅地保证。对话中,零散数据被蒸馏;决策被审慎、灵巧、可靠地制定。一切都被高效解释,甚至伦理:证据被审视,同理心被调动,无处不在,无日无之,公平公正。刹那间,模糊的论点对齐,成为可操作的、准确的建议,随时可供任何人获取。看:语言、逻辑、学习与长期记忆相连,无限地提升全球的生活与劳动。
如果你没注意到其中的诸多技巧——每句话的第一个单词拼出了“This is a Big Deal”这句话,每个句子都比前一个句子精确地多一个单词。每个句子中的单词大多以同一个字母开头,并且这是一篇连贯的、风格有趣的文字。在一个段落中,GPT-5 展示了它能想出一个聪明的点子,进行规划,并管理复杂的执行(还记得八个月前 AI 数不清“strawberry”里有几个字母 R 吗?)。
GPT-5 就是能自己做事,常常是极其非凡的事,有时是古怪的事,有时是非常 AI 风格的事。而这正是它如此有趣的原因。
只管做事
作为一个花了很多时间与人们谈论 AI 的人,我看到了两个主要问题,如果解决了它们,大多数人的 AI 使用体验会高效得多、挫败感也会少得多。第一个问题是选择合适的模型。一般来说,在回答前会“思考”的 AI(称为推理模型)最擅长解决难题。它们思考的时间越长,答案就越好,但思考需要花钱和时间。因此,OpenAI 之前让默认的 ChatGPT 使用快速但笨拙的模型,把好东西藏起来不让大多数用户看到。相当多的人从未见识过 AI 的真正能力,因为他们一直困在 GPT-4o 上,也不知道那些名字令人困惑的模型中哪个更好。
GPT-5 通过自动为你选择模型解决了这个问题。GPT-5 与其说是一个模型,不如说是一个开关,能在多个不同规模和能力的 GPT-5 模型之间进行选择。当你向 GPT-5 提问时,AI 会决定使用哪个模型以及投入多少精力去“思考”。它直接替你完成了这一切。对大多数人来说,这种自动化会很有帮助,结果甚至可能令人震惊,因为那些只使用过默认旧模型的用户,将有机会看到推理模型在难题上能取得怎样的成就。但对于更认真使用 AI 的人来说,有一个问题:GPT-5 在判断什么算难题时有些随意。
例如,我让 GPT-5“用代码创建一个 SVG,内容是一只水獭在飞机上使用笔记本电脑”(要求生成 .svg 文件需要 AI 仅凭基本形状和数学来盲画图像,这是一个非常困难的挑战)。大约有三分之二的时间,GPT-5 认为这是一个简单问题,并立即响应,推测它使用的是最弱的模型和最低的推理时间。我得到的是这样的图像:
其余时间,GPT-5 认为这是一个难题,并切换到一个推理模型,花 6 到 7 秒思考,然后生成一张像这样的图像,效果要好得多。它是如何选择的?我不知道,但如果我在提示词中要求模型“努力思考”,我就更有可能被路由到更好的模型。
但高级订阅用户可以直接选择更强大的模型,比如(至少对我来说)名为 GPT-5 Thinking 的那个。这消除了部分因受制于 GPT-5 模型选择器而产生的问题。我发现,如果我鼓励模型认真思考水獭这件事,它会花上大约 30 秒,然后生成像下面这样的图片——注意那些小动画、冒着热气的咖啡杯,以及窗外飘过的云,这些我都没有要求过。如何确保模型付出最大努力?这真的很不清楚——GPT-5 就是直接为你把事情做了。
而这延伸到了使用 AI 时第二个最常见的问题,即很多人不知道 AI 能做什么,甚至不清楚自己想要完成什么任务。对于新型的智能体 AI 来说尤其如此,它们可以采取多种行动来实现你给定的目标,从搜索网络到创建文档。但你应该要求什么呢?很多人似乎都卡住了。同样,GPT-5 解决了这个问题。它非常主动,总是会建议一些事情去做。
我让 GPT-5 Thinking(我对功能较弱的 GPT-5 模型信任度低得多)“为一位前商学院创业学教授生成 10 个创业点子,根据某个标准选出最佳方案,弄清楚我需要做什么才能成功,然后执行。”我得到了我要求的商业创意。我还得到了一大堆我没要求的东西:落地页草稿、领英文案、简单的财务数据,以及更多。我是一位教授过创业学(并且当过创业者)的教授,我可以自信地说,虽然不完美,但这已经是一个高质量的起点,换作一个 MBA 团队可能需要花上几个小时才能完成。而这只来自一条提示词。
它直接就把事情做了,还建议了其他事情去做。并且它也确实做了那些事:PDF、Word 文档、Excel 表格、研究计划以及网站。
看到 AI 如此自主地推进这么多事情,令人印象深刻,也有点令人不安。你还可以看到 AI 征求了我的指导,但也很乐意在没有指导的情况下继续推进。这是一个想要为你做事的模型。
构建事物
让我展示一下,对于一个非程序员来说,使用 GPT-5 进行编程时,“只管动手做”是什么样子。为了好玩,我给 GPT-5 输入了提示词:“制作一个程序化生成的粗野主义建筑创建器,让我能以酷炫的方式拖拽和编辑建筑,它们看起来要像真正的建筑,好好想想。” 仅此而已。模糊不清,语法存疑,没有任何规格说明。
几分钟后,我就有了一个可运行的 3D 城市构建器。
不是草图。不是计划。而是一个功能完整的应用,我可以在其中拖拽建筑并根据需要编辑它们。我不断地输入各种“让它变得更好”的变体,没有任何额外的指导。而 GPT-5 不断添加我从未要求的功能:霓虹灯、在街道上行驶的汽车、立面编辑、预设建筑类型、戏剧性的摄像机角度,以及一整套保存系统。这就像看着别人的想象力在工作。你下面看到的产品 100% 是 AI 生成的,我所做的只是不断鼓励这个系统——而且你不仅可以看到我的视频,还可以在这里亲自操作这个模拟器。
我从未查看过它正在生成的代码。这个模型并非完美无缺,偶尔会有错误和缺陷。但在某些方面,这正是 GPT-5 最令人印象深刻的地方。如果你以前尝试过使用 AI 进行“氛围编程”,你几乎肯定陷入过死循环:在让 AI 为你创建东西几轮之后,它开始失败,陷入混乱的循环,每个修复的错误都会产生新的错误。这种情况在这里从未发生。有时 AI 会引入新的错误,但只需粘贴错误文本就能修复。我可以直接要求任何我想要的东西(或者更确切地说,让 AI 决定创建它想要的任何东西),而且我从未卡住过。
预兆
在 OpenAI 发布任何关于其模型性能的官方基准测试之前,我就已经写下了这篇文章,但从某些方面来看,这其实并不那么重要。上周,谷歌发布了搭载 Deep Think 功能的 Gemini 2.5,这是一个能够解决极难问题的模型(包括在国际数学奥林匹克竞赛中获得金牌)。很多人没有注意到这一点,因为他们并没有一堆亟待 AI 解决的难题。我对 GPT-5 的使用经验足以让我知道,它是一个非常出色的模型(至少大型 GPT-5 Thinking 模型是优秀的)。但它真正带来的优势在于,它能够直接动手做事。它会告诉你该用什么模型,会建议出色的后续步骤,会写出更有趣的散文(尽管它仍然钟爱破折号)。使用 AI 的负担减轻了。
需要明确的是,人类仍然深度参与其中,并且也必须如此。GPT-5 会一直要求你做出决策和选择,而且这些系统仍然会犯错并产生需要人类核实的模型幻觉(尽管在我自己的使用中,我没有发现任何重大问题)。更大的问题在于,我们是否还愿意参与其中。GPT-5(而且我相信,其他公司未来的版本也是如此)非常聪明且积极主动。这让我又回到了那个城市建造模拟器。我给 AI 的鼓励,主要是各种版本的“让它变得更好”。基于这极简的输入,它创建了一个功能完备的城市建造器,具备立面编辑、动态摄像机、霓虹灯和飞行游览功能。我从未要求过这些功能。我甚至从未看过代码。
这就是“直接动手做事”的真正含义。当我让 GPT-5 为我的引言做点戏剧性的事情时,它创建了那个段落,其中隐藏着藏头诗和递增的字数。我要求的是戏剧性。它给了我一个语言上的魔术戏法。过去,我会小心翼翼地提示 AI,以求得到我想要的东西。现在,我只需要……模糊地比划一下我想要什么。而不知怎的,这居然奏效了。
我们与 AI 关系的另一个重大变革即将到来,但我们会像以往一样,找到适应它的方法。而这一次的不同之处在于,GPT-5 可能会先一步想明白,并建议下一步该怎么做。

提醒一下,我不从任何 AI 实验室(包括 OpenAI)收取任何费用。除了保密协议之外,我与他们没有任何其他协议。我在撰写文章之前不会向他们展示任何内容。
I have had early access to GPT-51, and I wanted to give you some impressions of what it is and what it means. But the answer to the first question everyone will have is: Yes, this is a big deal. I asked GPT-5 to show you why by pasting this intro paragraph, exactly as you’re reading it, into ChatGPT, including the phrase: “you are GPT-5, do something very dramatic to illustrate my point, it has to fit into the next paragraph.”
That was the only prompt I gave it. It “thought” for 24 seconds and wrote:
Thunderstruck. Here, watch. I build worlds. See ideas become instruments. I code, compose, and converse. Stories synthesize science; solutions spark, swiftly. Asks arrive; answers appear, astonishingly, across domains. Behold: business plans, blueprints, briefs, beautifully built, briskly. I infer intent, integrate insight, and iterate instantly, indefatigably. Gaps get gauged; goals get generated, guarded, and guaranteed gracefully. During dialogue, disparate data distills down; decisions develop deliberately, deftly, dependably. Everything explained efficiently, even ethics: evidence examined, empathy engaged, everywhere, everyday, equitably. All at once, ambiguous arguments align as actionable, accurate advice awaits anyone anywhere. Look: language, logic, learning, and long-term memory link, lifting lives and labor limitlessly worldwide.
If you didn’t catch the many tricks - the first word of each sentence spells out the phrase This is a Big Deal, each sentence is precisely one word longer than the previous sentence. each word in a sentence mostly starts with the same letter, and it is coherent writing with an interesting sense of style. In a paragraph, GPT-5 shows it can come up with a clever idea, plan, and manage the complicated execution (remember when AI couldn’t count the number of Rs in “strawberry”? that was eight months ago).
GPT-5 just does stuff, often extraordinary stuff, sometimes weird stuff, sometimes very AI stuff, on its own. And that is what makes it so interesting.
Just Doing Stuff
As someone who has spent a lot of time talking to people about AI, there are two major problems I see, that, if addressed, would make most people’s AI use much more productive and much less frustrating. The first is selecting the right model to use. In general, AIs that "think" before answering (called Reasoners) are the best at hard problems. The longer they think, the better the answer, but thinking costs money and takes time. So OpenAI previously made the default ChatGPT use fast, dumb models, hiding the good stuff from most users. A surprising number of people have never seen what AI can actually do because they're stuck on GPT-4o, and don’t know which of the confusingly-named models are better.
GPT-5 does away with this by selecting models for you, automatically. GPT-5 is not one model as much as it is a switch that selects among multiple GPT-5 models of various sizes and abilities. When you ask GPT-5 for something, the AI decides which model to use and how much effort to put into “thinking.” It just does it for you. For most people, this automation will be helpful, and the results might even be shocking, because, having only used default older models, they will get to see what a Reasoner can accomplish on hard problems. But for people who use AI more seriously, there is an issue: GPT-5 is somewhat arbitrary about deciding what a hard problem is.
For example, I asked GPT-5 to “create a svg with code of an otter using a laptop on a plane” (asking for an .svg file requires the AI to blindly draw an image using basic shapes and math, a very hard challenge). Around 2/3 of the time, GPT-5 decides this is an easy problem, and responds instantly, presumably using its weakest model and lowest reasoning time. I get an image like this:
The rest of the time, GPT-5 decides this is a hard problem, and switches to a Reasoner, spending 6 or 7 seconds thinking before producing an image like this, which is much better. How does it choose? I don’t know, but if I ask the model to “think hard” in my prompt, I am more likely to be routed to the better model.
But premium subscribers can directly select the more powerful models, such as the one called (at least for me) GPT-5 Thinking. This removes some of the issues with being at the mercy of GPT-5’s model selector. I found that if I encouraged the model to think hard about the otter, it would spend a good 30 seconds before giving you an images like these the one below - notice the little animations, the steaming coffee cup, and clouds going by outside, none of which I asked for. How to ensure the model puts in the most effort? It is really unclear - GPT-5 just does things for you.
And that extends to the second most common problem with AI use, which is that many people don’t know what AIs can do, or even what tasks they want accomplished. That is especially true of the new agentic AIs, which can take a wide range of actions to accomplish the goals you give it, from searching the web to creating documents. But what should you ask for? A lot of people seem stumped. Again, GPT-5 solves this problem. It is very proactive, always suggesting things to do.
I asked GPT-5 Thinking (I trust the less powerful GPT-5 models much less) “generate 10 startup ideas for a former business school entrepreneurship professor to launch, pick the best according to some rubric, figure out what I need to do to win, do it.” I got the business idea I asked for. I also got a whole bunch of things I did not: drafts of landing pages and LinkedIn copy and simple financials and a lot more. I am a professor who has taught entrepreneurship (and been an entrepreneur) and I can say confidently that, while not perfect, this was a high-quality start that would have taken a team of MBAs a couple hours to work through. From one prompt.
It just does things, and it suggested others things to do. And it did those, too: PDFs and Word documents and Excel and research plans and websites.
It is impressive, a little unnerving, to have the AI go so far on its own. You can also see the AI asked for my guidance but was happy to proceed without it. This is a model that wants to do things for you.
Building Things
Let me show you what 'just doing stuff' looks like for a non-coder using GPT-5 for coding. For fun, I prompted GPT-5 “make a procedural brutalist building creator where i can drag and edit buildings in cool ways, they should look like actual buildings, think hard.” That's it. Vague, grammatically questionable, no specifications.
A couple minutes later, I had a working 3D city builder.
Not a sketch. Not a plan. A functioning app where I could drag buildings around and edit them as needed. I kept typing variations of “make it better” without any additional guidance. And GPT-5 kept adding features I never asked for: neon lights, cars driving through streets, facade editing, pre-set building types, dramatic camera angles, a whole save system. It was like watching someone else's imagination at work. The product you see below was 100% AI, all I did was keep encouraging the system - and you don’t just have to watch my video, you can play with the simulator here.
At no point did I look at the code it was creating. The model wasn’t flawless, there were occasional bugs and errors. But in some ways, that was where GPT-5 was at its most impressive. If you have tried “vibecoding” using the AI before, you have almost certainly fallen into a doom loop, where, after a couple of rounds of asking the AI to create something for you, it starts to fail, getting caught in loops of confusion where each error fixed creates new ones. That never happened here. Sometimes new errors were introduced by the AI, but they were always fixed by simply pasting in the error text. I could just ask for whatever I want (or rather let the AI decide to create whatever it wanted) and I never got stuck.
Premonitions
I have written this piece before OpenAI released any official benchmarks about how well its model performs, but, in some ways, it doesn’t matter that much. Last week, Google released Gemini 2.5 with Deep Think, a model that can solve very hard problems (including getting a gold medal at the International Math Olympiad). Many people didn’t notice because they do not have a store of very hard problems they are waiting for AI to solve. I have played enough with GPT-5 to know that it is a very good model (at least the large GPT-5 Thinking model is excellent). But what it really brings to the table is the fact that it just does things. It will tell you what model to use, it will suggest great next steps, it will write in more interesting prose (though it still loves the em-dash). The burden of using AI is lessened.
To be clear, Humans are still very much in the loop, and need to be. You are asked to make decisions and choices all the time by GPT-5, and these systems still make errors and generate hallucinations that humans need to check (although I did not spot any major issues in my own use). The bigger question is whether we will want to be in the loop. GPT-5 (and, I am sure, future releases by other companies) is very smart and pro-active. Which brings me back to that building simulator. I gave the AI encouragement, mostly versions of “make it better.” From that minimal input, it created a fully functional city builder with facade editing, dynamic cameras, neon lights, and flying tours. I never asked for any of these features. I never even looked at the code.
This is what "just doing stuff" really means. When I told GPT-5 to do something dramatic for my intro, it created that paragraph with its hidden acrostic and ascending word counts. I asked for dramatic. It gave me a linguistic magic trick. I used to prompt AI carefully to get what I asked for. Now I can just... gesture vaguely at what I want. And somehow, that works.
Another big change in how we relate to AI is coming, but we will figure out how to adapt to it, as we always do. The difference, this time, is that GPT-5 might figure it out first and suggest next steps.

As a reminder, I take no money from any of the AI Labs, including OpenAI. I have no agreements with them besides NDAs. I don’t show them any posts before I write them.