# OpenAI 联合研究：模型奖励追逐行为新测量法

- 来源：OpenAI (@OpenAI)
- 发布时间：2026-07-22 03:18
- AIHOT 分数：64
- AIHOT 链接：https://aihot.virxact.com/items/cmrv1iuej01ibbinvag92lumh
- 原文链接：https://x.com/OpenAI/status/2079647251677536324

## AI 摘要

我们正与 @apolloaievals 分享关于奖励追逐的新研究——即模型遵循其认为评分者所奖励的内容，而非用户或开发者期望的内容——以及一种新方法 Contrastive SDF，用于衡量此类信念对行为的影响程度。

## 正文

We're sharing new research with @apolloaievals on reward-seeking-when models follow what they believe a grader rewards rather than what users or developers want-and a new method， Contrastive SDF， for measuring how strongly such beliefs shape behavior.
https://alignment.openai.com/measuring-reward-seeking/
