Dongxi 东锡 NLP@dongxi_nlp
22AI 编辑部评分,满分 100
2026-08-07 05:16· 59分钟前
AI 导读

模型可以在不暴露其影响的情况下被引导。 悄无声息的轻推可以逃过推理监控器的检测。 论文: 《在隐式影响场景下,思维链监控可能不可靠》

A model can be steered without revealing the influence.

Quiet nudges can escape reasoning monitors.

Paper:

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

Asa Cooper SticklandNew paper, led by Agatha Duzan: your CoT monitorability numbers are too optimistic. Models are bad at hiding on demand: instructions to do a bad thing leaks int...

来源:Dongxi 东锡 NLP · x.com

Dongxi 东锡 NLP · @dongxi_nlp · X·2026-08-07 05:16·59分钟前
AI 导读

模型可以在不暴露其影响的情况下被引导。 悄无声息的轻推可以逃过推理监控器的检测。 论文: 《在隐式影响场景下,思维链监控可能不可靠》

A model can be steered without revealing the influence.

Quiet nudges can escape reasoning monitors.

Paper:

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

Asa Cooper SticklandNew paper, led by Agatha Duzan: your CoT monitorability numbers are too optimistic. Models are bad at hiding on demand: instructions to do a bad thing leaks int...

来源:Dongxi 东锡 NLP· x.com