我刚刚在宾夕法尼亚大学上了一堂实验课,要求学生们在四天内从零开始创建一家初创公司。班上大部分学生就读于高级管理人员工商管理硕士(EMBA)项目,因此他们一边上课,一边还在各种大小公司担任医生、经理或领导者。几乎没有人写过代码。我向他们介绍了 Claude Code 和 Google Antigravity,他们需要用这些工具来构建一个可运行的原型。但仅有原型并不算一家初创公司,所以他们又使用 ChatGPT、Claude 和 Gemini 来加速创意生成、市场调研、竞争定位、路演以及财务建模等流程。我很想知道他们在这么短的时间内能走多远。结果发现,他们走得非常远。

我教授创业课程已有十五年之久,期间看过成千上万个创业想法(其中一些最终发展成了大公司),因此我很清楚一群聪明的MBA学生能取得怎样的成果。据我估计,我在短短几天内看到的成果,在通往真正创业公司的道路上,比我在AI出现前看到学生整个学期完成的成果要领先一个数量级。大多数原型不仅仅是示例界面,而是已经实现了核心功能。想法也比以往更加多样化和有趣。市场和客户分析也颇具洞察力。这确实令人印象深刻。这些还不是已经运作的初创公司,也不是完全可用的产品(少数例外),但它们从传统流程中节省了数月时间以及巨额资金和精力。还有一点:大多数早期初创公司需要调整方向,随着对市场需求和技术可能性的了解加深而改变路线。通过降低调整方向的成本,探索各种可能性变得容易得多,不会陷入僵局,甚至可以同时探索多个创业方向:你只需告诉AI你想要什么。
我希望我能说这些令人印象深刻的成果归功于我出色的教学,但我们实际上还没有一个完善的框架来指导如何使用所有这些工具,学生们很大程度上是自己摸索出来的。他们具备一些管理和专业领域知识这一点很有帮助,因为事实证明,成功的关键实际上就是上一段最后一点:告诉AI你想要什么。随着AI越来越能胜任需要人类花费数小时才能完成的任务,并且评估这些结果也变得越来越耗时,善于委派的价值也随之增加。但什么时候应该把任务委派给AI呢?
智能体工作方程
我们其实有答案,但情况有点复杂。需要考虑三个因素:第一,由于 AI 能力的“锯齿状边界”,你无法可靠地预判 AI 在复杂任务上哪些做得好、哪些做得差。第二,无论 AI 表现好坏,它都绝对快。它能在几分钟内完成人类需要数小时才能完成的工作。第三,它很便宜(相对于专业人员的薪资),而且你生成多个版本然后扔掉大部分,它也不介意。
这三个因素意味着,决定是否将任务委托给 AI 取决于三个变量:
人类基准时间:你自己完成该任务所需的时间
成功概率:AI 在单次尝试中产出符合你标准的输出的可能性
AI 处理时间:你发起请求、等待并评估 AI 输出所需的时间
一个有用的思维模型是,你在“完成整个任务”(人类基准时间)和“支付开销成本”(AI 处理时间)之间做权衡,可能需要多次重复,直到得到可接受的结果。成功概率越高,你需要支付 AI 处理时间的次数就越少,将任务交给 AI 就越有用。例如,考虑一个任务,你自己需要一小时完成,但 AI 几分钟就能搞定,不过检查答案需要三十分钟。在这种情况下,只有当成功概率非常高时,你才应该把工作交给 AI,否则你花在生成和检查草稿上的时间会比亲自完成还多。但如果人类基准时间是 10 小时,那么花几个小时与 AI 协作就是值得的,前提是 AI 能够胜任这份工作。

我们知道这个等式成立,因为今年夏天,OpenAI 发布了一篇关于人工智能与实际工作的重要论文——GDPval。我之前讨论过这篇论文,其关键在于,它让来自金融、医学、政府等多个领域的经验丰富的人类专家与最新的人工智能进行对决,并由另一组专家担任评委。人类专家平均需要七小时完成工作,因此,在这种情况下,这就是人类基线时间。人工智能处理时间则很有趣:人工智能完成任务只需几分钟,但专家实际检查工作结果需要一小时,当然,编写提示词也需要时间。至于成功概率,当 GDPval 首次发布时,评委在大多数情况下判定人类工作获胜;但随着 GPT-5.2 的发布,天平发生了倾斜。GPT-5.2 Thinking 和 Pro 模型平均有 72% 的概率与人类专家打成平手或超越他们。

现在,假设成功概率为 72%,评估时间为一小时,我们可以计算出在一个七小时的任务上你能节省多少小时。如果你尝试每个任务时都花时间向人工智能编写提示词,花一小时评估其回答,如果人工智能回答不佳再亲自动手,那么你平均可以节省 3 小时。人工智能失败的任务会花费更长时间(因为你浪费了时间编写提示词和审查!),但人工智能成功的任务则会快得多。不过,我们可以利用管理学技巧,让这个等式对我们更有利!
委托即新的提示词
我们可以通过三种方式提高委托 AI 的成功概率并降低 AI 处理时间,从而让委托 AI 更有价值。我们可以给出更好的指令,设定明确的目标,让 AI 能够以更高的成功率执行。我们可以提升评估和反馈的能力,从而减少让 AI 做对事情所需的尝试次数。我们还可以在不花费太多时间的情况下,更轻松地评估 AI 在某项任务上的表现好坏。所有这些因素都会因领域专业知识而得到改善——专家知道该给出什么指令,他们能更好地发现何时出现问题,也更擅长纠正问题。
如果你不需要特定的东西,AI 模型在自行解决问题方面已经变得极其强大。例如,我发现 Claude Code 能够仅凭一条提示词就生成一整个 1980 年代风格的冒险游戏,这条提示词是:“创作一款完全原创的老派 Sierra 风格冒险游戏,采用类似 EGA 的图形。你应该使用你的图像智能体来生成图片,并给我一个解析器。让所有谜题都既有趣又可解。完成整个游戏(游玩时间应在 10-15 分钟),不要问任何问题。让它变得惊艳且令人愉悦。”仅此而已,AI 制作了一切,包括美术。通过最后两条提示词,它测试了游戏并进行了部署。你可以亲自体验:enchanted-lighthouse-game.netlify.app
这确实令人惊叹,但这种惊叹感之所以被放大,是因为我不需要任何特定的东西,只需要一个 AI 可以自由即兴创作的冒险游戏。但真正的工作和真正的委托,意味着你心中有一个特定的输出目标,而这正是事情可能变得棘手的地方。你如何向 AI 传达你的意图,让它执行你想要的,使其既能运用“判断力”来解决问题,同时又能给出你期望的输出?
这个问题早在人工智能出现之前就已存在,并且普遍到每个领域都发明了自己的文书工作来解决它。软件开发者编写产品需求文档。电影导演交接分镜表。建筑师创建设计意图文件。海军陆战队使用五段式命令(态势、任务、执行、行政、指挥)。咨询顾问通过详细的交付物规格来界定项目范围。所有这些文档在智能体工作的新时代里,作为 AI 提示词都表现得非常出色(而且 AI 一次能处理很多页的指令)。之所以能用这么多格式来指导 AI,是因为它们本质上都是同一件事:试图把一个人脑子里的想法转化为另一个人的行动。
当你审视优秀授权文档的实际内容时,会发现它们惊人地一致:我们要完成什么,以及为什么?授权的边界在哪里?“完成”是什么样子的?我需要哪些具体输出?我需要哪些阶段性输出来跟进你的进展?在告诉我你完成之前,你应该检查什么?如果这些内容都明确指定了,那么 AI 和人类一样,更有可能把工作做好。
而在弄清楚如何向 AI 下达这些指令的过程中,你实际上是在重新发明管理。
管理 AI 智能体
我觉得有趣的是,观察到一些主要 AI 实验室里最知名的软件开发者指出,他们的工作正从以编程为主转变为以管理 AI 智能体为主。编程一直有着非常严谨的结构,输出结果可以清晰验证(代码要么能运行,要么不能),因此它成为 AI 工具最早成熟的领域之一,也是第一个感受到这种变化的职业。但这不会是最后一个。
作为一名商学院教授,我认为许多人已经具备或能够学会与 AI 智能体协作所需的技能——这些本质上就是管理学 101 的基础能力。如果你能清晰说明需求、给予有效反馈、设计评估工作的方法,你就能与智能体良好协作。从很多方面来看,至少在你擅长的领域,这比设计精巧的提示词来完成任务要容易得多,因为它更像与人共事。与此同时,管理学始终建立在资源稀缺的假设之上:你之所以委派任务,是因为无法事必躬亲,也因为人才既有限又昂贵。AI 改变了这个等式。如今,“人才”变得丰富且廉价。真正稀缺的是知道该提出什么要求。
这正是我的学生们表现出色的原因。他们并非 AI 专家。但他们花了多年时间学习如何在自己的专业领域界定问题、明确交付物,以及识别财务模型或医疗报告中的异常。他们从课程和工作中积累了来之不易的框架,而这些框架最终成为了他们的提示词。那些常被轻视的“软技能”,恰恰成了最硬核的能力。
当每个人都成为管理者,手下拥有一支不知疲倦的智能体大军时,我无法确切预知工作的形态。但我猜想,那些能够脱颖而出的人,将是那些懂得什么是“好”的标准——并且能够清晰表达,以至于连 AI 都能据此交付成果的人。我的学生们在四天内就领悟了这一点。不是因为他们天生就懂 AI,而是因为他们早已懂得如何管理。原来,他们所受的全部训练,恰好是在为这一刻做准备。
I just taught an experimental class at the University of Pennsylvania where I challenged students to create a startup from scratch in four days. Most of the people in the class were in the executive MBA program, so they were taking classes while also working as doctors, managers, or leaders in a variety of large and small companies. Few had ever coded. I introduced them to Claude Code and Google Antigravity, which they needed to use to build a working prototype. But a prototype alone is not a startup, so they used ChatGPT, Claude, and Gemini to accelerate the idea generation, market research, competitive positioning, pitching, and financial modelling processes. I was curious how far they could get in such a short time. It turns out they got very far.

I’ve been teaching entrepreneurship for a decade and a half, and I've seen thousands of startup ideas (some of which turned into large companies) so I have a good sense of the expectations for what a class of smart MBA students can accomplish. I would estimate that what I saw in a couple of days was an order of magnitude further along the path to a real startup than I had seen out of students working over a full semester before AI. Most of the prototypes were not just sample screens but actually had a core feature working. Ideas were far more diverse and interesting than usual. Market and customer analyses were insightful. It was really impressive. These were not yet working startups nor were they fully operational products (with a couple exceptions) — but they had shaved months and huge amounts of money and effort from the traditional process. And there was something else: most early startups need to pivot, changing direction as they learn more about what the market wants and what is technically possible. By lowering the costs of pivoting, it was much easier to explore the possibilities without being locked in or even explore multiple startups at once: you just tell the AI what you want.
I wish I could say this impressive output was the result of my brilliant teaching, but we don’t really have a great framework yet for how to use all these tools, the students largely figured it out on their own. It helped that they had some management and subject matter expertise because it turns out that the key to success was actually the last bit of the previous paragraph: telling the AI what you want. As AIs are increasingly capable of tasks that would take a human hours to do, and as evaluating those results becomes increasingly time consuming, the value of being good at delegation increases. But when should you delegate to AI?
The Equation of Agentic Work
We actually have an answer, but it is a bit complicated. Consider three factors: First, because of the Jagged Frontier of AI ability, you don’t reliably know what the AI will be good or bad at on complex tasks. Second, whether the AI is good or bad, it is definitely fast. It produces work in minutes that would take many hours for a human to do. Third, it is cheap (relative to professional wages), and it doesn’t mind if you generate multiple versions and throw most of them away.
These three factors mean that deciding to delegate to AI depends on three variables:
Human Baseline Time: how long the task would take you to do yourself
Probability of Success: how likely the AI is to produce an output that meets your bar on a given attempt
AI Process Time: how long it takes you to request, wait for, and evaluate an AI output
A useful mental model is that you’re trading off “doing the whole task” (Human Baseline Time) against “paying the overhead cost” (AI Process Time), possibly multiple times until you get something acceptable. The higher Probability of Success is, the fewer times you have to pay AI Process Time, and the more useful it is to turn things over to the AI. For example, consider a task that takes you an hour to do, but the AI can do it in minutes, though checking the answer takes thirty minutes. In that case, you should only give the work to the AI if Probability of Success is very high, otherwise you’ll spend more time generating and checking drafts than just doing it yourself. If the Human Baseline Time is 10 hours, though, it could be worth several hours of working with the AI, assuming that the AI can be made to do a competent job.

We know this equation works because this past summer, OpenAI released one of the more important papers on AI and real work, GDPval. I have discussed it before, but the key was that it pitted experienced human experts in diverse fields from finance to medicine to government against the latest AIs, with another set of experts working as judges. It took experts seven hours on average to do the work, so, in this case, that is the Human Baseline Time. The AI Process Time was interesting: the AI took only minutes for tasks, but it required an hour for experts to actually check the work, and, of course, prompts take time to write as well. As for Probability of Success, when GDPval first came out, judges gave human work the win the majority of the time, but, with the release of GPT-5.2, the balance shifted. GPT-5.2 Thinking and Pro models tied or beat human experts an average of 72% of the time.

We can now calculate how many hours you would save on a seven-hour task, assuming that 72% probability of success and an hour of evaluation. If you tried every task by taking the time to prompt the AI, evaluating the answer for an hour, and then doing it yourself if the AI answer was bad, you would save 3 hours on average. Tasks the AI failed on would take longer (you wasted time prompting and reviewing!) but tasks the AI succeeded on would be much faster. But we can change the equation even more in our favor using techniques from management!
Delegation as the new prompting
There are three things we can do to make delegating to AI more worthwhile by increasing the Probability of Success and lowering AI Process Time. We can give better instructions, setting clear goals that the AI can execute on with a higher chance of succeeding. We can get better at evaluation and feedback, so we need to make fewer attempts to get the AI to do the right thing. And we can make it easier to evaluate whether the AI is good or bad at a task without spending as much time. All of these factors are improved by subject matter expertise — an expert knows what instructions to give, they can better see when something goes wrong, and they are better at correcting it.
If you don’t need something specific, AI models have become incredibly capable of figuring out how to solve problems themselves. For example, I found Claude Code was able to generate an entire 1980s style adventure game with one prompt to "create an entirely original old-school Sierra style adventure game with EGA-like graphics. You should use your image agent to generate images and give me a parser. Make all puzzles interesting and solvable. Finish the game (it should take 10-15 minutes to play), don’t ask any questions. make it amazing and delightful." That’s it, the AI made everything, including the art. With two final prompts it tested the game and deployed it. You can play it yourself: enchanted-lighthouse-game.netlify.app
This is genuinely amazing, but that amazement is amplified because I didn’t need anything specific, just an adventure game that the AI was free to improvise. But real work, and real delegation, means that you have a specific output in mind, and that is where things can get tricky. How do you communicate your intention to the AI to execute on what you want, so it can use “judgement” to solve problems while still giving you the output you desire?
This problem existed long before AI and is so universal that every field has invented their own paperwork to solve it. Software developers write Product Requirements Documents. Film directors hand off shot lists. Architects create design intent documents. The Marines use Five Paragraph Orders (situation, mission, execution, administration, command). Consultants scope engagements with detailed deliverable specs. All of these documents work remarkably well as AI prompts for this new world of agentic work (and the AI can handle many pages of instructions at a time). The reason you can use so many formats to instruct AI is that all of these are really the same thing: attempts to get what’s in one person’s head into someone else’s actions.
When you look at what actually goes into good delegation documentation, it’s remarkably consistent: What are we trying to accomplish, and why? Where are the limits of the delegated authority? What does “done” look like? What specific outputs do I need? What interim outputs do I need to follow your progress? And what should you check before telling me you’re finished? If these are well-specified, the AI, like humans, is far more likely to do a good job.
And in figuring out how to give these instructions to the AI, it turns out you are basically reinventing management.
Managing Agents
I find it interesting to watch as some of the most well-known software developers at the major AI labs note how their jobs are changing from mostly programming to mostly management of AI agents. Coding has always had a very organized structure, with clearly verifiable outputs (the code either works or it doesn’t) so it has been one of the first areas where AI tools have matured, and thus the first profession to feel this change. It isn’t the last.
As a business school professor, I think many people have the skills they need, or can learn them, in order to work with AI agents - they are management 101 skills. If you can explain what you need, give effective feedback, and design ways of evaluating work, you are going to be able to work with agents. In many ways, at least in your area of expertise, it is much easier than trying to design clever prompts to help you get work done, as it is more like working with people. At the same time, management has always assumed scarcity: you delegate because you can’t do everything yourself, and because talent is limited and expensive. AI changes the equation. Now the “talent” is abundant and cheap. What’s scarce is knowing what to ask for.
This is why my students did so well. They weren’t AI experts. But they’d spent years learning how to scope problems in their fields of expertise, define deliverables, and recognize when a financial model or medical report was off. They had hard-earned frameworks from classes and jobs, and those frameworks became their prompts. The skills that are so often dismissed as “soft” turned out to be the hard ones.
I don’t know exactly what work looks like when everyone is a manager with an army of tireless agents. But I suspect the people who thrive will be the ones who know what good looks like — and can explain it clearly enough that even an AI can deliver it. My students figured this out in four days. Not because they were AI natives, but because they already knew how to manage. All that training, it turns out, was accidentally preparing them for exactly this moment.