Very interesting new work from Microsoft Research.
(bookmark it)
They show that a 4B model can be tuned to out-negotiate the GPT-5 family of models.
The dispositions that make an assistant pleasant make it a poor delegate. They show that a friendly frontier model volunteers its principal's private information and concedes at the first sign of resistance.
SocialRL trains social reasoning directly in a 4B model across six principal-driven domains including negotiation, job interviews and marketplace haggling. After training, 78% of buyer openings anchor below target against 3% untrained.
Cascade RL and multi-teacher distillation consolidate the specialists into one 4B at 0.627 average utility, above GPT-5.1 at 0.619 and GPT-5.2 at 0.613.
Paper: https://arxiv.org/abs/2608.13787
Track more trending AI papers in our academy: https://academy.dair.ai/