Noam Brown · @polynoamial · X·2026-07-03 13:04·60天前
AI 导读

大多数AI智能体评估将能力归结为一个分数。但该数字隐藏了一个关键选择:智能体被允许使用的计算量。新工作展示了为什么这很重要。Noam Brown称赞这是优秀工作。

Noam Brown@polynoamial
30AI 编辑部评分,满分 100
2026-07-03 13:04· 60天前
AI 导读

大多数AI智能体评估将能力归结为一个分数。但该数字隐藏了一个关键选择:智能体被允许使用的计算量。新工作展示了为什么这很重要。Noam Brown称赞这是优秀工作。

Excellent work from @AISecurityInst investigating the impact of test-time compute budgets for frontier AI model evaluations. They make the case even more convincingly than I could!

AI Security InstituteMost AI agent evaluations boil capability down to one score. But that number hides a key choice: how much compute the agent was allowed to use. New work from ou...