Rohan Paul@rohanpaul_ai
43AI 编辑部评分,满分 100
2026-08-03 04:20· 26分钟前
跳到正文
AI 摘要

新论文《Would You Walk to the Car Wash?》提出 SaliTrap 基准,测试 12 个模型在 1,145 个提示中的常识推理,发现 LLM 存在“显著偏差”:显式数字和程序会压过未言明的物理前提。最佳模型仅 54.8% 避开陷阱,12 个中 8 个低于 30%;移除干扰后,四个代表性模型的正确率回升至 86.9%-91.8%。

LLMs can know a task is impossible and still optimize it anyway.

Ask whether to walk or drive to a car wash 50 meters away, and some models focus on distance while missing that the car itself must reach the wash.

The failure here is not missing knowledge, but failing to use it when explicit details dominate the prompt.

The paper calls this salience bias: explicit numbers and procedures overpower unstated physical prerequisites.

SaliTrap tests 1,145 such prompts across four trap types and 12 models.

Even the best evaluated model avoided the trap in only 54.8% of queries, while eight of twelve stayed below 30%.

More distractors made this worse: trap avoidance fell as numerical density rose, while models increasingly began calculating before noticing the contradiction.

Awareness was not enough either.

Among trap-aware responses, GLM-5.1 and Kimi-K2 still complied 86.2% and 81.8% of the time.

The strongest diagnosis comes from removing the bait: context-free probes recovered 86.9% to 91.8% of previously sycophantic cases across four representative models.

So many failures reflect knowledge that is present but suppressed by task framing.

For agent evaluation, premise checking should be tested as a behavioral control, not assumed from a model's general reasoning score.

  • arxiv. org/abs/2607.28478

Title: "Would You Walk to the Car Wash? Revealing the Salience Bias of LLMs in Commonsense Reasoning"

Rohan Paul · @rohanpaul_ai · X·2026-08-03 04:20·26分钟前
在 X 看原推· x.com
AI 摘要

新论文《Would You Walk to the Car Wash?》提出 SaliTrap 基准,测试 12 个模型在 1,145 个提示中的常识推理,发现 LLM 存在“显著偏差”:显式数字和程序会压过未言明的物理前提。最佳模型仅 54.8% 避开陷阱,12 个中 8 个低于 30%;移除干扰后,四个代表性模型的正确率回升至 86.9%-91.8%。

LLMs can know a task is impossible and still optimize it anyway.

Ask whether to walk or drive to a car wash 50 meters away, and some models focus on distance while missing that the car itself must reach the wash.

The failure here is not missing knowledge, but failing to use it when explicit details dominate the prompt.

The paper calls this salience bias: explicit numbers and procedures overpower unstated physical prerequisites.

SaliTrap tests 1,145 such prompts across four trap types and 12 models.

Even the best evaluated model avoided the trap in only 54.8% of queries, while eight of twelve stayed below 30%.

More distractors made this worse: trap avoidance fell as numerical density rose, while models increasingly began calculating before noticing the contradiction.

Awareness was not enough either.

Among trap-aware responses, GLM-5.1 and Kimi-K2 still complied 86.2% and 81.8% of the time.

The strongest diagnosis comes from removing the bait: context-free probes recovered 86.9% to 91.8% of previously sycophantic cases across four representative models.

So many failures reflect knowledge that is present but suppressed by task framing.

For agent evaluation, premise checking should be tested as a behavioral control, not assumed from a model's general reasoning score.

  • arxiv. org/abs/2607.28478

Title: "Would You Walk to the Car Wash? Revealing the Salience Bias of LLMs in Commonsense Reasoning"

在 X 查看原推x.com