微软4B模型谈判能力超越GPT-5系列

DAIR.AI · @dair_ai · X·2026-08-17 23:50·7天前
AI 导读

微软研究院新研究表明,4B模型经SocialRL训练后可在谈判中胜过GPT-5系列模型。训练后78%的买家开场报价低于目标价,而未训练模型仅3%。Cascade RL与多教师蒸馏将专家整合为单一4B模型,平均效用0.627,高于GPT-5.1的0.619和GPT-5.2的0.613。

DAIR.AI@dair_ai
47AI 编辑部评分,满分 100

微软4B模型谈判能力超越GPT-5系列

2026-08-17 23:50· 7天前
AI 导读

微软研究院新研究表明,4B模型经SocialRL训练后可在谈判中胜过GPT-5系列模型。训练后78%的买家开场报价低于目标价,而未训练模型仅3%。Cascade RL与多教师蒸馏将专家整合为单一4B模型,平均效用0.627,高于GPT-5.1的0.619和GPT-5.2的0.613。

Very interesting new work from Microsoft Research.

(bookmark it)

They show that a 4B model can be tuned to out-negotiate the GPT-5 family of models.

The dispositions that make an assistant pleasant make it a poor delegate. They show that a friendly frontier model volunteers its principal's private information and concedes at the first sign of resistance.

SocialRL trains social reasoning directly in a 4B model across six principal-driven domains including negotiation, job interviews and marketplace haggling. After training, 78% of buyer openings anchor below target against 3% untrained.

Cascade RL and multi-teacher distillation consolidate the specialists into one 4B at 0.627 average utility, above GPT-5.1 at 0.619 and GPT-5.2 at 0.613.

Paper: https://arxiv.org/abs/2608.13787

Track more trending AI papers in our academy: https://academy.dair.ai/