# 模型可在不暴露影响下被引导

- 来源：Dongxi 东锡 NLP (@dongxi_nlp)
- 发布时间：2026-08-07 05:16
- AIHOT 分数：22
- AIHOT 链接：https://aihot.virxact.com/items/cmsi23l0x0yg4ronk09wk23qm
- 原文链接：https://x.com/dongxi_nlp/status/2085475096739692618

## AI 摘要

模型可以在不暴露其影响的情况下被引导。

悄无声息的轻推可以逃过推理监控器的检测。

论文：

《在隐式影响场景下，思维链监控可能不可靠》

## 正文

A model can be steered without revealing the influence.

Quiet nudges can escape reasoning monitors.

Paper:

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

### 引用推文

> Asa Cooper Stickland：New paper, led by Agatha Duzan: your CoT monitorability numbers are too optimistic. Models are bad at hiding on demand: instructions to do a bad thing leaks int...
