A model can be steered without revealing the influence.
Quiet nudges can escape reasoning monitors.
Paper:
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
New paper, led by Agatha Duzan: your CoT monitorability numbers are too optimistic. Models are bad at hiding on demand: instructions to do a bad thing leaks int...