我听到各地首席财务官都在问一个简单的问题:我们如何从人工智能投入中获得更多价值?
多年来,市场一直通过采用率来衡量软件的成功:购买的席位、活跃的用户、续订的许可证。理解人工智能的价值需要一种更有力的衡量标准:完成的工作量。
首席财务官及其他企业领导者面临的基本经济问题是:人工智能完成的工作价值,其增长速度是否快于其生产成本的增长速度。
要回答这个问题,需要比“每 token 成本”这样的指标看得更深入。成本较低的模型可能 token 更便宜,但要获得出色的结果可能需要更多尝试、更多时间或更多人工审核。能力更强的模型可能 token 更昂贵,但能一次性完成同样的任务。关键在于产生一个成功结果的完整成本,并对照该结果所创造的价值来衡量。
人工智能时代的终极记分卡可以看作是“每美元的有用智能”。这个指标回答了四个关键问题:
- 人工智能是否在完成有意义的工作?
- 每项成功任务的成本是多少?
- 人们能否信赖其结果?
- 随着使用量的增长,每投入一美元的人工智能是否创造了更多价值?
1. 完成了多少有用工作?
从工作本身开始。
人工智能帮助解决了多少客户问题?帮助交付了多少代码变更?审查了多少份合同?为人们节省了多少时间?有多少决策因在正确时机获得了正确的上下文而得到改善?
当 token 转化为人们可以使用的实际工作时,它们就创造了价值。随着模型能力越来越强,它们可以承担更长、更复杂的任务:保持上下文、进行多步骤推理、跨工具协作,并在过程中不断调整适应。
最好的起点是选择一个工作流程。定义“完成”的含义,并在工作发生的系统中衡量该结果。
对于支持团队来说,“完成”可能意味着一个客户问题得到解决。对于工程团队来说,可能意味着一个通过测试的代码变更。对于法务团队来说,可能意味着一份合同得到准确且及时的审查。
设想一个财务团队正在为预测评审做准备。在最终决策做出之前,大量工作已经展开:查找最新预测、将数据导入 Excel 或 Sheets、识别变化、核对标签页、重建幻灯片,并检查所有数字是否完全吻合。
ChatGPT Work 可以承担这一过程中的大部分工作,让团队有更多时间专注于真正重要的问题:发生了什么变化?为什么?我们下一步该怎么做?
这就是实践中每一美元所能带来的实用智能。更多工作得以更快完成,而人们可以将更多时间用于运用判断力、创造力和专业知识。
2. 一个成功的任务实际花费是多少?
接下来的问题是,高质量完成这项工作需要多少成本。
AI 任务的差异很大。一个快速回答可能只需要很少的计算量。而编码、研究或财务工作流可能涉及更深入的推理、工具使用以及大量操作。这些更复杂的任务可能需要更多计算资源,但它们也能创造更大的价值。
在模型层面,每个成功任务的成本取决于价格、所使用的计算量以及得出正确结果的可能性。对于企业而言,总成本还包括员工时间、人工审核、重试和返工。
计算方法很简单:
- 加上完成工作的总成本。
- 统计达到所需质量标准的任务数量。
- 用总成本除以成功任务的数量。
这就是为什么最低的每 token 价格并不总能带来最低的每结果成本。即使对于常规请求,一个前沿模型也可能提供最佳价值,如果它能一次性给出正确答案,从而减少重试、延迟、审核和总计算量。
分层模型系列为客户提供了更多优化这一公式的途径。我们上周发布的 GPT‑5.6 有三个层级:Sol 是我们的旗舰模型;Terra 在性能与成本之间取得平衡;Luna 是我们最快且最经济的模型。
这些层级提供了有用的起点。整个任务的经济性最终应决定合适的模型。客户可能会为快速、高吞吐量的工作流使用 Luna,为需要更高深度的工作使用 Terra,或者在更强的推理能力能以更少尝试次数带来最佳结果时使用 Sol。
我们训练了 GPT‑5.6,让每个 token 都能产出更多有用成果。在 Artificial Analysis 编程智能体指数(Artificial Analysis Coding Agent Index)上,开启最大推理的 GPT‑5.6 Sol 创下了新的最优水平,同时比另一款领先模型少用了 54% 的输出 token。下图展示了这一对比。
DeepSWE v1.1:在长周期工程任务中,GPT‑5.6 Sol 达到了 72.7% 的新高,高于 Claude Fable 5 的 69.9%,且预估 API 成本降低了 36.2%。
在整个 GPT‑5.6 系列中,目标是一致的:每美元产出更多成功成果。更高的效率让现有任务变得更经济。更强的能力则让全新类型的工作成为可能。
每一代新模型都应在这两方面有所提升。客户应当能够完成更有价值的工作,同时每项任务的完成成本持续下降。
3. AI 正确完成工作的频率有多高?
第三个衡量指标是可靠性。
AI 的采用通常会分阶段深化。首先,AI 辅助起草。然后,它开始在工具和数据之间查找上下文并进行推理。随着时间的推移,它开始采取行动、处理异常并完成工作流程,而人类则在需要时提供判断和控制。
每一步都创造更多价值,也对系统提出更高要求。
可靠性具有直接的经济价值。当结果准确、来源可靠、前后一致且能恰当升级处理时,人们花在审查、纠正和重复工作上的时间就会减少。成功完成的任务成本更低,组织也更有信心在更重要的流程中使用 AI。
团队可以通过追踪以下三种结果来具体衡量这一点:
- 可直接使用:结果在交付时达到了质量标准。
- 需要修正:结果需要再次尝试或人工编辑。
- 需要升级处理:需要人工介入并完成工作。
这些衡量指标比单纯的模型准确率更能说明问题。它们能显示 AI 是否真正减少了完成项目所需的工作量。
可靠性还需要明确的边界。在 AI 从起草阶段转向采取行动之前,组织应定义:
- 系统可以访问哪些数据。
- 系统可以使用或更改哪些系统。
- 何时需要人工审查或批准某个操作。
安全、保障、隐私和控制构成了深度使用的基础。人们需要了解系统如何运作、他们的数据如何处理,以及系统的行为如何被管控。
ChatGPT Work 建立在 ChatGPT Enterprise 的安全、保障、合规和工作区管理基础之上。这使得组织能够在保持适当监督的同时,为 AI 提供更多上下文并接入更有价值的工作流程。
能力带来首次使用,可靠性则让 AI 成为完成工作的组成部分。
4. 随着使用量增长,每一美元 AI 投入能否完成更多工作?
最后一个问题是,随着规模扩大,经济效益是否也会提升。
企业可以通过长期追踪同一工作流程来衡量这一点。统计达到质量标准的任务数量、完成这些任务的总成本,以及每项成功任务的成本。如果完成的工作量增长快于总成本,同时质量保持不变或有所提升,那么每一美元 AI 投入就产生了更多价值。
算力处于这一等式的核心位置。
算力驱动着研究以及 AI 完成的每一项任务。它决定了产品质量、速度、可靠性、可用性和成本。训练算力构建未来能力,推理算力则交付当下的实用价值。两者都应转化为更好的客户成果。
更优的模型、更高效的推理、专用硬件、更高的利用率、更智能的路由以及更强的产品设计,都能提升算力的回报。每一代基础设施都有助于训练出能力更强的模型。而更好的算法、硬件和软件,则能更高效地服务于这些模型。
客户从人性化的角度体验这些改进:更准确的答案、更快的响应、更少的修正、更可靠的产品,以及完成所需工作更低的成本。
这些收益会不断累积。更好的基础设施加速研究,研究催生能力更强、效率更高的模型,更好的模型改进产品,更优的产品推动采用、学习和收入增长。这种增长反过来支持对下一代研究、算力、部署和安全领域的持续投入。
OpenAI 通过一个统一的智能平台将这些要素整合在一起。用户通过 ChatGPT 和 ChatGPT Work 使用它。开发者通过 Codex 和 API 基于它进行构建。企业将其部署到实际工作发生的系统中。
当某一层得到改进时,每一个产品和客户都能从中受益。
AI 时代的记分卡
综合来看,这四项衡量指标告诉我们,每单位成本所能获得的有用智能是否在提升。
有用产出告诉我们 AI 能产生什么。每项成功任务的成本告诉我们达成结果需要付出什么。可靠性告诉我们人们可以放心使用多少工作成果。规模化价值告诉我们,随着时间的推移,每一美元和每一单位算力是否实现了更多成效。
目标是让 AI 帮助人们从事更有意义的工作,做出更明智的决策,并将更多时间投入到那些需要独特的人类判断力和创造力的工作环节中。
我们的职责是让这个等式随着每一代模型变得更好:更强大的模型、更快更可靠的结果,以及为客户所需工作降低的成本。
这就是 AI 如何随着时间的推移,对更多人和组织变得更有用。
The question I hear from CFOs everywhere is simple: how do we get more value from our AI spend?
For years, the market measured the success of software through adoption: seats purchased, users active, licenses renewed. Understanding the value of AI demands a more powerful measure: work accomplished.
The basic economic question facing CFOs and other business leaders is whether the value of the work AI completes grows faster than the cost of producing it.
Answering that question requires looking more deeply than a metric such as cost per token. A lower-cost model may have cheaper tokens, but getting great results may require more attempts, more time, or more human review. A more capable model may have more expensive tokens, but complete the same task in one pass. What matters is the full cost of producing a successful outcome, measured against the value that outcome creates.
The ultimate scorecard for the age of AI could be looked at as “Useful Intelligence per Dollar.” This metric answers four key questions:
- Is AI completing work that matters?
- What does each successful task cost?
- Can people depend on the result?
- Does each AI dollar produce more value as usage grows?
1. How much useful work gets done?
Start with the work itself.
How many customer issues did AI help resolve? How many code changes did it help ship? How many contracts did it review? How much time did it give back to people? How many decisions improved because the right context was available at the right moment?
Tokens create value when they transform into work people can use. As models become more capable, they can take on longer and more complex tasks: maintaining context, reasoning through multiple steps, working across tools, and adapting as they go.
The best place to begin is with one workflow. Define what “done” means and measure that outcome in the system where the work happens.
For a support team, “done” might mean a customer issue resolved. For an engineering team, it might mean a code change that passes its tests. For a legal team, it might mean a contract reviewed accurately and on time.
Consider a finance team preparing for a forecast review. Much of the work happens before a final decision is made: finding the latest forecast, moving data into Excel or Sheets, identifying changes, reconciling tabs, rebuilding slides, and checking that everything adds up perfectly.
ChatGPT Work can take on much of that process, giving the team more time to focus on the questions that matter: What changed? Why? What should we do next?
That is useful intelligence per dollar in practice. More work gets completed, faster, while people spend more of their time applying judgment, creativity, and expertise.
2. What does a successful task actually cost?
The next question is what it costs to complete that work well.
AI tasks vary widely. A quick answer may require little compute. A coding, research, or financial workflow may involve deeper reasoning, tool use, and many actions. Those more complex tasks can require more compute, but they can create much more value.
At the model level, cost per successful task depends on price, the amount of compute used, and the likelihood of reaching the right result. For a business, the full cost also includes employee time, human review, retries, and rework.
The calculation is straightforward:
- Add the full cost of completing the work.
- Count the tasks that met the required quality bar.
- Divide the full cost by the number of successful tasks.
This is why the lowest price per token does not always produce the lowest cost per outcome. A frontier model may deliver the best value even for a routine request if it produces the right answer in one pass, reducing retries, latency, review, and total compute.
A tiered model family gives customers more ways to optimize this equation. GPT‑5.6, which we released last week, has three tiers: Sol is our flagship; Terra balances performance and cost; Luna is our fastest and most affordable model.
These tiers provide useful starting points. The economics of the full task should ultimately determine the right model. A customer might use Luna for a fast, high-volume workflow, Terra for work requiring greater depth, or Sol when stronger reasoning delivers the best result with fewer attempts.
We trained GPT‑5.6 to get more useful work from every token. On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol with max reasoning set a new state of the art while using 54% fewer output tokens than another leading model. The chart below illustrates the comparison.
DeepSWE v1.1: Long-horizon engineering tasks; GPT‑5.6 Sol reaches a new high of 72.7%, above Claude Fable 5’s 69.9%, at 36.2% lower estimated API cost.
Across the GPT‑5.6 family, the goal is the same: more successful work per dollar. Greater efficiency makes existing tasks more affordable. Greater capability makes entirely new kinds of work possible.
Each new model generation should improve both sides of that equation. Customers should be able to accomplish more valuable work while, at the same time, the cost of completing each task continues to fall.
3. How often does AI get the work right?
The third measure is dependability.
AI adoption tends to deepen in stages. First, AI helps draft. Then it finds context and reasons across tools and data. Over time, it begins taking action, handling exceptions, and completing workflows, with people providing judgment and control where needed.
Each step creates more value and asks more of the system.
Dependability has direct economic value. When results are accurate, well-sourced, consistent, and escalated appropriately, people spend less time reviewing, correcting, and repeating the work. Successful tasks cost less, and organizations gain the confidence to use AI in more important workflows.
Teams can make this concrete by tracking three outcomes:
- Ready to use: The result met the quality bar as delivered.
- Needs correction: The result required another attempt or human edits.
- Needs escalation: A person needed to step in and finish the work.
These measures tell a richer story than model accuracy alone. They show whether AI is genuinely reducing the work involved in completing the project.
Dependability also requires clear boundaries. Before AI moves from drafting to taking action, organizations should define:
- What data the system can access.
- What systems it can use or change.
- When a person should review or approve an action.
Safety, security, privacy, and control create the foundation for deeper use. People need to understand how the system behaves, how their data is handled, and how its actions are governed.
ChatGPT Work builds on the security, privacy, compliance, and workspace-management foundation of ChatGPT Enterprise. This allows organizations to give AI more context and access to more valuable workflows while maintaining appropriate oversight.
Capability earns first use. Dependability makes AI part of how work gets done.
4. Does each AI dollar buy more work as usage grows?
The final question is whether the economics improve at scale.
Companies can measure this by following the same workflow over time. Track how many tasks met the quality bar, the total cost of completing them, and the cost per successful task. If completed work grows faster than total cost while quality holds or improves, each AI dollar is producing more value.
Compute sits at the center of this equation.
Compute powers research and every task that AI completes. It shapes product quality, speed, dependability, availability, and cost. Training compute builds future capability. Inference compute delivers useful work today. Both should translate into better outcomes for customers.
Better models, more efficient inference, purpose-built hardware, higher utilization, smarter routing, and stronger product design all improve the return on compute. Each generation of infrastructure helps train more capable models. Better algorithms, hardware, and software then help serve those models more efficiently.
Customers experience those improvements in human terms: better answers, faster results, fewer corrections, more dependable products, and a lower cost for the work they need done.
The gains compound. Better infrastructure accelerates research. Research produces more capable and efficient models. Better models improve products. Better products drive adoption, learning, and revenue. That growth supports continued investment in the next generation of research, compute, deployment, and safety.
OpenAI brings these pieces together through one shared intelligence platform. People use it through ChatGPT and ChatGPT Work. Developers build with it through Codex and the API. Enterprises deploy it into the systems where work happens.
When one layer improves, every product and customer can benefit.
A scorecard for the AI age
Taken together, these four measures tell us whether useful intelligence per dollar is improving.
Useful work tells us what AI produces. Cost per successful task tells us what it takes to reach the outcome. Dependability tells us how much of the work people can confidently use. Value at scale tells us whether each dollar, and each unit of compute, accomplish more over time.
The goal is AI that helps people do more meaningful work, make better decisions, and spend more time on the parts of their jobs that require distinctly human judgment and creativity.
Our job is to make that equation better with every generation: more capable models, faster and more dependable results, and lower costs for the work customers need done.
That is how AI becomes more useful to more people and organizations over time.