# EoBench：措辞变化如何影响LLM对错误主张的接受度

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-07-22 12:29
- AIHOT 分数：47
- AIHOT 链接：https://aihot.virxact.com/items/cmrvlw57d03n3bihbnmc53pf8
- 原文链接：https://x.com/rohanpaul_ai/status/2079785990504206702

## AI 摘要

EoBench用约6.6万条以19种风格（形式、证据、确定性、语气）编写的错误主张，测试LLM是否信任用户新主张还是依赖训练事实。在Gemma、Llama和Qwen的18个模型上，命令式、儿童化用语、正式语言和权威主张说服力最强，而弱主张和反事实说服力最弱。更大模型和指令微调通常能减少对错误上下文的跟随，提示措辞会悄然改变模型答案。

## 正文

LLMs can accept the same false claim differently depending on its tone， certainty， and grammatical form.

Small wording changes can make LLMs accept false claims， while larger and instruction-tuned models resist them more.

Models must decide whether to trust a user's new claim or rely on facts stored during training.

EoBench tests this choice with about 66K false claims written in 19 styles across form， evidence， certainty， and tone.

The team evaluated 18 Gemma， Llama， and Qwen models， then kept cases where each model already knew the correct fact.

Commands， child-directed wording， formal language， and authority claims persuaded models most， while weak claims and counterfactuals persuaded them least.

Across Llama and Gemma， larger models followed false context less often， and instruction tuning usually reduced that behavior.

The finding shows that prompt wording can quietly change model answers， so evaluations and product safeguards must test linguistic framing directly.

---

- arxiv. org/abs/2607.18232

Title： "It's Not What You Say， It's How You Say It： Evaluating LLM Responses to Expressions of Belief"
