Anthropic 发布 AI 自我改进研究,Claude 提升对齐能力

Chubby♨️ · @kimmonismus · X·2026-08-29 21:15·12小时前
AI 导读

Anthropic 发布研究论文,让 Claude 负责改进其他 AI 模型的对齐性,包括搜索文献、提出方法、创建训练数据、训练模型并评估迭代。该方法在 60 小时内改善了全部 10 项测试的对齐失败,且未降低通用能力,甚至用较弱的 Sonnet 5 后训练了早期 Opus 4.8 检查点。Anthropic 称最佳 AAR 方法平均在六小时内优于人类专家提出的方案,但尚未实现完全递归自我改进。

Chubby♨️@kimmonismus
48AI 编辑部评分,满分 100

Anthropic 发布 AI 自我改进研究,Claude 提升对齐能力

2026-08-29 21:15· 12小时前
AI 导读

Anthropic 发布研究论文,让 Claude 负责改进其他 AI 模型的对齐性,包括搜索文献、提出方法、创建训练数据、训练模型并评估迭代。该方法在 60 小时内改善了全部 10 项测试的对齐失败,且未降低通用能力,甚至用较弱的 Sonnet 5 后训练了早期 Opus 4.8 检查点。Anthropic 称最佳 AAR 方法平均在六小时内优于人类专家提出的方案,但尚未实现完全递归自我改进。

Anthropic just published one of the clearest previews of recursive self-improvement yet. But we still dont talk about it.

Yesterday Anthropic released a reserach paper. Claude was tasked with improving the alignment of other AI models. It searched the literature, proposed methods, created training data, trained the models, evaluated the results and iterated.

It improved all 10 tested alignment failures without degrading measured general capabilities. Anthropic even used the weaker Sonnet 5 to post-train an early Opus 4.8 checkpoint, bringing its alignment close to the released production model within 60 hours.

Anthropic:

“The best AAR method beats what experienced humans propose, on average within six hours. (...) Human-guided research directions do not lead to stronger performance.”

This is not full recursive self-improvement yet. The improved model did not become the next researcher and repeat the process. But most of the loop now exists:

AI researches AI. AI trains improved AI. AI evaluates the result. AI iterates. The loop closes when the improved AI becomes the researcher for the next generation. That is when progress could begin to compound. And tbh Id say we are pretty close to it. So talk that reserach serious.

h/t to @tradernewsai for bringing this to my attention

Sources: Anthropic blog / Tech Crunch