Rohan Paul@rohanpaul_ai
50AI 编辑部评分,满分 100
2026-08-04 15:32· 9分钟前
跳到正文
AI 摘要

最新研究显示,最强AI模型能在多轮博弈中维持连贯欺骗策略,而非仅产生单一谎言。Kimi K2.5通过早期帮助对手建立信誉、嫁祸无辜玩家,最终成功让秘密队友当选,全程难以被识别;GPT-5.4同样保持隐蔽。较弱模型则因决策与叙述不一致而暴露。研究基于ParliamentBench社交推理游戏,论文见arxiv.org/abs/2607.28146。

The strongest AI models did not just lie well. They stayed believable round after round.

The strongest models did far more than produce one convincing lie.

They kept their story, votes, and strategy aligned across many rounds, so others continued to trust them while they quietly moved the game toward a hidden goal.

In one match, Kimi K2.5 helped the opposing side early to build credibility, later blamed an innocent player for a bad outcome, and then used that trust to get its secret teammate elected.

Deception here is not one false sentence; it is a plan carried through memory, timing, persuasion, and action.

Weaker models exposed themselves when their decisions stopped matching their story, while Kimi K2.5 and GPT-5.4 stayed difficult to identify throughout the game.

  • arxiv. org/abs/2607.28146

Title: "Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game"

Rohan Paul · @rohanpaul_ai · X·2026-08-04 15:32·9分钟前
在 X 看原推· x.com(在新标签页打开)
AI 摘要

最新研究显示,最强AI模型能在多轮博弈中维持连贯欺骗策略,而非仅产生单一谎言。Kimi K2.5通过早期帮助对手建立信誉、嫁祸无辜玩家,最终成功让秘密队友当选,全程难以被识别;GPT-5.4同样保持隐蔽。较弱模型则因决策与叙述不一致而暴露。研究基于ParliamentBench社交推理游戏,论文见arxiv.org/abs/2607.28146。

The strongest AI models did not just lie well. They stayed believable round after round.

The strongest models did far more than produce one convincing lie.

They kept their story, votes, and strategy aligned across many rounds, so others continued to trust them while they quietly moved the game toward a hidden goal.

In one match, Kimi K2.5 helped the opposing side early to build credibility, later blamed an innocent player for a bad outcome, and then used that trust to get its secret teammate elected.

Deception here is not one false sentence; it is a plan carried through memory, timing, persuasion, and action.

Weaker models exposed themselves when their decisions stopped matching their story, while Kimi K2.5 and GPT-5.4 stayed difficult to identify throughout the game.

  • arxiv. org/abs/2607.28146

Title: "Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game"

在 X 查看原推x.com(在新标签页打开)