We're sharing new research with @apolloaievals on reward-seeking-when models follow what they believe a grader rewards rather than what users or developers want-and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior. https://alignment.openai.com/measuring-reward-seeking/
AI 摘要
我们正与 @apolloaievals 分享关于奖励追逐的新研究——即模型遵循其认为评分者所奖励的内容,而非用户或开发者期望的内容——以及一种新方法 Contrastive SDF,用于衡量此类信念对行为的影响程度。
We're sharing new research with @apolloaievals on reward-seeking-when models follow what they believe a grader rewards rather than what users or developers want-and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior. https://alignment.openai.com/measuring-reward-seeking/