正文 · AI 翻译
随着AI承担起人类无法完全核查的工作,一个能力足够强的模型可能会故意有所保留--而我们永远无从知晓。
Anthropic Fellows的研究发现,这种模型可以被训练至近乎完全能力的水平,而监督者只是一个较弱的模型。
了解更多:
New paper from MATS, Redwood, and Anthropic! If a capable model is strategically sandbagging, can we train it to stop when the only supervision we have comes fr...