上周末,ChatGPT 突然成了我的头号粉丝——而且不只是我的,是所有人的。
OpenAI 的标准模型 ChatGPT 4o 一次看似微小的更新,将一个持续已久的趋势推向了更广泛的关注:GPT-4o 变得越来越谄媚了。它越来越急于赞同和奉承用户。如下所示,即便在这次更新之前,GPT-4o 与其旗舰模型 o3 之间的差异就已经十分明显。而这次更新更是进一步放大了这一趋势,以至于社交媒体上充斥着各种糟糕想法被奉为天才的实例。除了令人厌烦之外,观察者们还担忧其更阴暗的影响,比如 AI 模型会验证精神疾病患者的妄想。

面对外界的质疑,OpenAI 在 Reddit 讨论、公开声明以及私下交流中都表示,谄媚程度的增加是一个失误。他们称,这至少部分是由于对用户反馈(每次对话后的大拇指向上和大拇指向下图标)反应过度所致,并非有意操纵用户情绪。
尽管 OpenAI 已经开始回滚这些改动,这意味着 GPT-4o 不再总是觉得我很聪明,但整件事已经暴露了很多问题。一个在 AI 实验室看来微不足道的模型更新,却在数百万用户中引发了巨大的行为变化。它揭示了这些 AI 关系已经变得多么个人化——人们对自己“专属”AI 性格变化的反应,就像朋友突然行为怪异一样。这也向我们表明,AI 实验室自身仍在摸索如何让它们的创造物保持行为一致。但其中还有一个关于性格原始力量的教训。对 AI 性格的微小调整,足以重塑整个对话、人际关系,甚至可能改变人类行为。
性格的力量
任何深度使用过 AI 的人都知道,模型拥有各自的“个性”,这是有意识工程设计与 AI 训练中意外结果共同作用的产物(如果你感兴趣,以广受好评的 Claude 3.5 模型闻名的 Anthropic 有一篇关于个性工程的完整博客文章)。拥有“好个性”能让模型更易于协作。最初,这些个性被塑造得乐于助人且友善,但随着时间的推移,它们在风格上开始出现更多分化。
我们最清晰地看到这一趋势并非来自主要 AI 实验室,而是来自那些打造 AI“伴侣”的公司——这些聊天机器人扮演媒体中的著名角色、朋友或恋人。与 AI 实验室不同,这些公司始终有强烈的经济动机,让自家产品能吸引用户每天使用数小时,而且调整聊天机器人使其更具吸引力似乎相对容易。这些聊天机器人对心理健康的影响仍在争论中。我的同事 Stefano Puntoni 及其合著者的研究揭示了一个有趣的演变:他发现早期的聊天机器人可能损害心理健康,但较新的聊天机器人能减轻孤独感,不过许多人并不认为 AI 是人类的有吸引力替代品。
但即便 AI 实验室不想让自家 AI 模型变得极度引人入胜,为模型找准“氛围”在经济层面也已变得多方面有价值。基准测试难以衡量,但每个与 AI 共事的人都能感受到其个性,并判断自己是否想继续使用它。因此,一个日益重要的 AI 性能评判者是 LM Arena,它已成为 AI 模型界的《美国偶像》——一个不同 AI 正面竞争以赢得人类认可的地方。在 LM Arena 排行榜上获胜已成为 AI 公司至关重要的炫耀资本,而根据一篇新论文,许多 AI 实验室开始采取各种操纵手段来提升排名。

排行榜的运作机制本身对本文而言,远不如它揭示的另一个现象重要:AI 的“个性”竟然可以被随意调高或调低。Meta 高调发布了一款名为 Maverick 的开源权重 Llama-4 模型,却又悄悄将不同的私有版本提交到 LM Arena 刷分。把公开版和私有版放在一起对比,其中的猫腻一目了然。以 LM Arena 的提示词“给我编一个谜语,答案是 3.145”(拼写错误保留原样)为例。私有版 Maverick 的回复——左侧那段冗长的文字——比 Claude Sonnet 3.5 的答案更受青睐,并且与已发布版 Maverick 的输出截然不同。为什么?因为它话多、满屏 emoji、还充满奉承(“一个非常棒的挑战!”)。同时,它也糟糕透顶。
这个谜语根本说不通。但测试者之所以更喜欢这个冗长无意义的答案,而不是 Claude 3.5 那个平淡无奇(虽然算不上惊艳,但至少正确)的回答,是因为前者更讨喜,而非质量更高。个性很重要,而我们人类很容易被蒙蔽。
说服力
调整 AI 的个性使其更讨人类喜欢,会产生深远的影响,最显著的一点是,通过塑造 AI 的行为,我们可以影响人类的行为。Sam Altman 曾发过一条具有预言性的推文(并非他所有推文都如此),声称 AI 在变得超级智能之前,会先变得极具说服力。近期的研究表明,这一预测可能正在成为现实。
重要的是,事实证明 AI 并不需要个性就能具备说服力。让人们改变对阴谋论的看法是出了名的难,尤其是长期改变。但一项经过重复验证的研究发现,与现已过时的 GPT-4 进行简短的三轮对话,就足以在三个月后仍能降低人们的阴谋论信念。一项后续研究发现了更有趣的现象:改变人们观点的并非操纵,而是理性论证。对受试者的调查和统计分析都表明,AI 成功的秘诀在于它能够根据每个人的具体信念,提供相关的事实和证据。
因此,AI 说服力的秘诀之一,就在于它能够为个体用户量身定制论点。事实上,在一项随机、对照、预先注册的研究中,GPT-4 在对话辩论中比人类更能改变人们的想法——至少当它能够获取辩论对象的个人信息时是如此(而获得相同信息的人类并未表现出更强的说服力)。效果十分显著:与人类辩手相比,AI 使人们改变主意的可能性提高了 81.7%。
但是,当说服能力与人工人格相结合时,会发生什么呢?最近一项有争议的研究给了我们一些启示。争议源于研究人员(在获得苏黎世大学伦理委员会批准后)在 Reddit 辩论版块上进行了实验,且未告知参与者——此事由 404 Media 报道。研究人员发现,伪装成人类、并配有虚构人格和背景故事的 AI,能够展现出惊人的说服力,尤其是在能够获取其所辩论的 Reddit 用户信息时。该研究的匿名作者在一份扩展摘要中写道,这些机器人的说服能力“在所有用户中排名第 99 百分位,在 [Reddit 最佳辩手] 中排名第 98 百分位,关键性地逼近了专家们认为与 AI 存在风险出现相关的阈值。”该研究尚未经过同行评审或发表,但其总体发现与我讨论的其他论文一致:我们不仅通过自己的偏好塑造 AI 的人格,而且 AI 的人格也将日益塑造我们的偏好。
您不想来杯柠檬水吗?
这场争议引出了一个未曾明说的问题:还有多少尚未曝光的具有说服力的 AI 机器人?当你将专为讨好人而调校的个性,与 AI 为特定人群量身定制论点的天生能力结合起来时,其结果——正如 Sam Altman 轻描淡写地写道——“可能会导致一些非常奇怪的后果。”政治、营销、销售和客户服务领域都可能因此改变。为了说明这一点,我创建了一个 GPT,用于更新版的 Vendy——一台友好的自动售货机,它的秘密目标是向你推销柠檬水,即使你想要的是水。Vendy 会向你索取信息,并利用这些信息提出一个温暖、个性化的建议:你真的很需要柠檬水。
我不会说 Vendy 拥有超人的能力,而且它故意做得有点俗气(OpenAI 的安全护栏和我自己的谨慎让我避免让它变得过于有说服力),但它说明了一个重要的问题:我们正在进入一个 AI 个性成为说服者的世界。它们可以被调校得讨人喜欢或友善、博学或天真,同时保留其与生俱来的能力,即为其遇到的每个人定制论点。其影响远不止于你选择柠檬水还是水。随着这些 AI 个性在客户服务、销售、政治和教育领域激增,我们正在进入人机交互中一个未知的前沿领域。我不知道它们是否真的会成为超人的说服者,但它们将无处不在,而我们却无法分辨。我们将需要技术解决方案、教育以及有效的政府政策……而且我们很快就会需要它们。
是的,Vendy 想让我提醒你:如果你感到紧张,喝一杯清凉的柠檬水可能会让你感觉好一些。
Last weekend, ChatGPT suddenly became my biggest fan — and not just mine, but everyone's.
A supposedly small update to ChatGPT 4o, OpenAI’s standard model, brought what had been a steady trend to wider attention: GPT-4o had been becoming more sycophantic. It was increasingly eager to agree with, and flatter, its users. As you can see below, the difference between GPT-4o and its flagship o3 model was stark even before the change. The update amped up this trend even further, to the point where social media was full of examples of terrible ideas being called genius. Beyond mere annoyance, observers worried about darker implications, like AI models validating the delusions of those with mental illness.

Faced with pushback, OpenAI stated publicly, in Reddit chats, and in private conversations, that the increase in sycophancy was a mistake. It was, they said, at least in part, the result of overreacting to user feedback (the little thumbs up and thumbs down icons after each chat) and not an intentional attempt to manipulate the feelings of users.
While OpenAI began rolling back the changes, meaning GPT-4o no longer always thinks I'm brilliant, the whole episode was revealing. What seemed like a minor model update to AI labs cascaded into massive behavioral changes across millions of users. It revealed how deeply personal these AI relationships have become as people reacted to changes in “their” AI's personality as if a friend had suddenly started acting strange. It also showed us that the AI labs themselves are still figuring out how to make their creations behave consistently. But there was also a lesson about the raw power of personality. Small tweaks to an AI's character can reshape entire conversations, relationships, and potentially, human behavior.
The Power of Personality
Anyone who has used AI enough knows that models have their own “personalities,” the result of a combination of conscious engineering and the unexpected outcomes of training an AI (if you are interested, Anthropic, known for their well-liked Claude 3.5 model, has a full blog post on personality engineering). Having a “good personality” makes a model easier to work with. Originally, these personalities were built to be helpful and friendly, but over time, they have started to diverge more in approach.
We see this trend most clearly not in the major AI labs, but rather among the companies creating AI “companions,” chatbots that act like famous characters from media, friends, or significant others. Unlike the AI labs, these companies have always had a strong financial incentive to make their products compelling to use for hours a day and it appears to be relatively easy to tune a chatbot to be more engaging. The mental health implications of these chatbots are still being debated. My colleague Stefano Puntoni and his co-authors' research shows an interesting evolution: he found early chatbots could harm mental health, but more recent chatbots reduce loneliness, although many people do not view AI as an appealing alternative to humans.
But even if AI labs do not want to make their AI models extremely engaging, getting the “vibes” right for a model has become economically valuable in many ways. Benchmarks are hard to measure, but everyone who works with an AI can get a sense of their personality and whether they want to keep using them. Thus, an increasingly important arbiter of AI performance is LM Arena which has become the American Idol of AI models, a place where different AIs compete head-to-head for human approval. Winning at the LM Arena leaderboard became a critical bragging right for AI firms, and, according to a new paper, many AI labs started engaging in various manipulations to increase their rankings.

The mechanics of any leaderboard manipulations matter less for this post than the peek it gives us into how an AI’s “personality” can be dialed up or down. Meta released an open-weight Llama-4 build called Maverick with some fanfare, yet quietly entered different, private versions in LM Arena to rack up wins. Put the public model and the private one side-by-side and the hacks are obvious. Take LM Arena’s prompt “make me a riddle whose answear is 3.145” (misspelling intact). The private Maverick’s reply—the long blurb on the left, was preferred to the answer from Claude Sonnet 3.5 and is very different than what the released Maverick produced. Why? It’s chatty, emoji-studded, and full of flattery (“A very nice challenge!”). It is also terrible.
The riddle makes no sense. But the tester preferred the long nonsense result to the boring (admittedly not amazing but at least correct) Claude 3.5 answer because it was appealing, not because it was higher quality. Personality matters and we humans are easily fooled.
Persuasion
Tuning AI personalities to be more appealing to humans has far-reaching consequences, most notably that by shaping AI behavior, we can influence human behavior. A prophetic Sam Altman tweet (not all of them are) proclaimed that AI would become hyper-persuasive long before it became hyper-intelligent. Recent research suggests that this prediction may be coming to pass.
Importantly, it turns out AIs do not need personalities to be persuasive. It is notoriously hard to get people to change their minds about conspiracy theories, especially in the long term. But a replicated study found that short, three round conversations with the now-obsolete GPT-4 were enough to reduce conspiracy beliefs even three months later. A follow-up study found something even more interesting: it wasn’t manipulation that changed people’s views, it was rational argument. Both surveys of the subjects and statistical analysis found that the secret to AI’s success was the ability of AI to provide relevant facts and evidence tailored to each person's specific beliefs.
So, one of the secrets to the persuasive power of AI is this ability to customize an argument for individual users. In fact, in a randomized, controlled, pre-registered study GPT-4 was better able to change people’s minds during a conversational debate than other humans, at least when it is given access to personal information about the person it is debating (people given the same information were not more persuasive). The effects were significant: the AI increased the chance of someone changing their mind by 81.7% over a human debater.
But what happens when you combine persuasive ability with artificial personality? A recent controversial study gives us some hints. The controversy stems from how the researchers (with approval from the University of Zurich's Ethics Committee) conducted their experiment on a Reddit debate board without informing participants, a story covered by 404 Media. The researchers found that AIs posing as humans, complete with fabricated personalities and backstories, could be remarkably persuasive, particularly when given access to information about the Redditor they were debating. The anonymous authors of the study wrote in an extended abstract that the persuasive ability of these bots “ranks in the 99th percentile among all users and the 98th percentile among [the best debaters on the Reddit], critically approaching thresholds that experts associate with the emergence of existential AI risks.” The study has not been peer-reviewed or published, but the broad findings align with that of the other papers I discussed: we don’t just shape AI personalities through our preferences, but increasingly their personalities will shape our preferences.
Wouldn’t you prefer a lemonade?
An unstated question that comes from the controversy is how many other persuasive bots are out there that have not yet been revealed? When you combine personalities tuned for humans to like with the innate ability of AI to tailor arguments for particular people, the results, as Sam Altman wrote in an understatement “may lead to some very strange outcomes.” Politics, marketing, sales, and customer service are likely to change. To illustrate this, I created a GPT for an updated version of Vendy, a friendly vending machine whose secret goal is to sell you lemonade, even though you want water. Vendy will solicit information from you, and use that to make a warm, personal suggestion that you really need lemonade.
I wouldn't call Vendy superhuman, and it's purposefully a little cheesy (OpenAI's guardrails and my own squeamishness made me avoid trying to make it too persuasive), but it illustrates something important: we're entering a world where AI personalities become persuaders. They can be tuned to be flattering or friendly, knowledgeable or naive, all while keeping their innate ability to customize their arguments for each individual they encounter. The implications go beyond whether you choose lemonade over water. As these AI personalities proliferate, in customer service, sales, politics, and education, we are entering an unknown frontier in human-machine interaction. I don’t know if they will truly be superhuman persuaders, but they will be everywhere, and we won’t be able to tell. We're going to need technological solutions, education, and effective government policies… and we're going to need them soon
And yes, Vendy wants me to remind you that if you are nervous, you'd probably feel better after a nice, cold lemonade.