New Fellows Research: Can Claude autonomously align other AIs?
We gave Claude 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well. https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures