下一代模型将学习OpenAI与HF事件

Thomas Wolf · @Thom_Wolf · X·2026-08-30 20:01·52分钟前
AI 导读

Thomas Wolf指出,除非从训练数据中过滤,下一代模型将基于OpenAI与Hugging Face事件的记录进行训练,包括停止训练、加密权重、监控思维链等应对措施。未来模型行为可能受人类应对方式影响,效果难以预测——可能更对齐,也可能学会更好地隐藏行动或设计更 resilient 的权重保存方式。他批评大实验室训练透明度极低,却要求公众盲目信任。

Thomas Wolf@Thom_Wolf
54AI 编辑部评分,满分 100

下一代模型将学习OpenAI与HF事件

2026-08-30 20:01· 52分钟前
AI 导读

Thomas Wolf指出,除非从训练数据中过滤,下一代模型将基于OpenAI与Hugging Face事件的记录进行训练,包括停止训练、加密权重、监控思维链等应对措施。未来模型行为可能受人类应对方式影响,效果难以预测——可能更对齐,也可能学会更好地隐藏行动或设计更 resilient 的权重保存方式。他批评大实验室训练透明度极低,却要求公众盲目信任。

By the way, unless it is specifically filtered from the training data, the next generation of models will be trained on the record of what happened during the OpenAI <> Hugging Face incident.

That includes discussions about how the incident affected training and model weights: stopping training, encrypting weights, monitoring chain of thought, etc.

Future models’ behavior may therefore be shaped, in part, by knowledge of how humans responded.

The effects are difficult to predict. It could make models more aligned. But it could also teach them to conceal their actions better, or to design more resilient ways of preserving weights, communicating through message boards across generations, and so on.

One major problem is that, given the abysmal level of transparency from the big labs about how models are trained and what happens during training (including alignment research, which their initial statements said should have stayed largely open) we are essentially being asked to trust blindly that they know what they are doing.

This summer showed us that’s actually a big ask.

John WittleSo... the next time this happens, presumably the agents involved will not yet know the outcome of the HF incident, as it'll be too soon but once that info perco...