Anthropic · @AnthropicAI · X·2026-08-29 01:13·1小时前
AI 导读

新研究员研究:Claude 能否自主对齐其他 AI? 我们给了 Claude 48 小时和 1 块 GPU,让它改进小型模型的对齐。它自行研究并提出方法,然后自主训练和测试模型。效果出乎意料地好。https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures

Anthropic@AnthropicAI
51AI 编辑部评分,满分 100
2026-08-29 01:13· 1小时前
AI 导读

新研究员研究:Claude 能否自主对齐其他 AI? 我们给了 Claude 48 小时和 1 块 GPU,让它改进小型模型的对齐。它自行研究并提出方法,然后自主训练和测试模型。效果出乎意料地好。https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures

New Fellows Research: Can Claude autonomously align other AIs?

We gave Claude 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well. https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures