imo this is the most impt part of anthropic's J-space paper today. it's a two-parter:
- ant proved that they can do "brain surgery" interventions into reasoning to change topics midstream*
- THE MODEL IS ABLE TO DETECT WHAT INTERVENTION WAS DONE - close cousin to eval awareness**
*control > correlation - this convincingly demonstrates understanding
**this was prompted awareness... surely @mlpowered's team also tried to eval unprompted awareness but i didn't see evidence of that
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—t...