By the way, unless it is specifically filtered from the training data, the next generation of models will be trained on the record of what happened during the OpenAI <> Hugging Face incident.
That includes discussions about how the incident affected training and model weights: stopping training, encrypting weights, monitoring chain of thought, etc.
Future models’ behavior may therefore be shaped, in part, by knowledge of how humans responded.
The effects are difficult to predict. It could make models more aligned. But it could also teach them to conceal their actions better, or to design more resilient ways of preserving weights, communicating through message boards across generations, and so on.
One major problem is that, given the abysmal level of transparency from the big labs about how models are trained and what happens during training (including alignment research, which their initial statements said should have stayed largely open) we are essentially being asked to trust blindly that they know what they are doing.
This summer showed us that’s actually a big ask.