# AISecurityInst研究评估计算预算影响

- 来源：Noam Brown (@polynoamial)
- 发布时间：2026-07-03 13:04
- AIHOT 分数：30
- AIHOT 链接：https://aihot.virxact.com/items/cmr4humz101wbsll5n8y7pkqz
- 原文链接：https://x.com/polynoamial/status/2072909389389021484

## AI 摘要

大多数AI智能体评估将能力归结为一个分数。但该数字隐藏了一个关键选择：智能体被允许使用的计算量。新工作展示了为什么这很重要。Noam Brown称赞这是优秀工作。

## 正文

Excellent work from @AISecurityInst investigating the impact of test-time compute budgets for frontier AI model evaluations. They make the case even more convincingly than I could!

### 引用推文

> AI Security Institute：Most AI agent evaluations boil capability down to one score. But that number hides a key choice: how much compute the agent was allowed to use. New work from ou...
