LLMs can accept the same false claim differently depending on its tone, certainty, and grammatical form.
Small wording changes can make LLMs accept false claims, while larger and instruction-tuned models resist them more.
Models must decide whether to trust a user's new claim or rely on facts stored during training.
EoBench tests this choice with about 66K false claims written in 19 styles across form, evidence, certainty, and tone.
The team evaluated 18 Gemma, Llama, and Qwen models, then kept cases where each model already knew the correct fact.
Commands, child-directed wording, formal language, and authority claims persuaded models most, while weak claims and counterfactuals persuaded them least.
Across Llama and Gemma, larger models followed false context less often, and instruction tuning usually reduced that behavior.
The finding shows that prompt wording can quietly change model answers, so evaluations and product safeguards must test linguistic framing directly.
---
- arxiv. org/abs/2607.18232
Title: "It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief"