简而言之:《经济学人》让 25 个前沿 AI 模型参与了世界价值观调查。模型所属实验室对价值观的预测能力,弱于训练与对齐选择:Gemini 与 Qwen 是近邻,GPT-4o 与 DeepSeek R1 近乎双胞胎,而 DeepSeek R1 与 DeepSeek V4 Flash 却形同陌路。世界观在代码生成中不可见,但在商业分析、预测、招聘和政策工作中,它是一个活跃的输入变量。
《经济学人》让 25 个前沿 AI 模型完成了世界价值观调查——这份问卷自 1981 年以来一直用于绘制 100 个国家的道德观念图谱。在这张 2x2 的坐标图中,有两个轴:第一轴是从传统(宗教)到世俗;第二轴是从生存(关注集体基本需求)到自我表达与个人主义。
大多数模型位于地图的自我表达半区,考虑到训练数据,这并不令人意外。
令人惊讶的是,模型之间的差异很大。Gemini 3.1 Flash Lite 与 Qwen 3.6 Flash 作为近邻,位于自我表达程度最高的区域。
GPT-4o 与 DeepSeek R1 近乎双胞胎,一个在旧金山训练,一个在杭州训练。
DeepSeek R1 与 DeepSeek V4 Flash 来自同一个实验室,却在世俗/传统轴线上处于完全相反的两端。
共享的训练数据和相似的标注者解释了“双胞胎”现象,而不同的后训练选择则解释了“陌生人”现象。Common Crawl 数据集中 46% 是英文内容,因此模型模仿的基础声音是一个受过大学教育的美国网民。Anthropic 随后将 Claude 对齐到《联合国人权宣言》中的原则,而这份宣言本身就是一个自由主义色彩的文件。
Grok 独树一帜,是一个传统的独立派。
这种差异改变了采购清单。如今,企业模型的每一份招标书都会评估价格、延迟、上下文窗口和基准分数。世界观并不在清单上。它应该被纳入吗?
对于代码生成、SQL、日志解析和图像分类来说,这没问题。计算机程序没有政治立场。
一旦模型被用于特定市场的商业决策,其世界观就成为一个活跃的输入变量。营销文案、用户行为预测以及客户支持的语气,都必须与目标人群的价值观相匹配。
AI 的世界观从未被视为 AI 采购的一部分,但对于某些用例来说,它可能有必要成为一个考量因素。
-
《经济学人》:AI 模型的价值取向与大多数人的价值观存在显著差异 ↩︎
-
Common Crawl 统计数据:语言分布情况 ↩︎
-
Anthropic:Claude 的宪法 ↩︎
In short : The Economist scored 25 frontier AI models on the World Values Survey. Lab of origin is a weaker predictor than training & alignment choices : Gemini & Qwen are neighbors, GPT-4o & DeepSeek R1 are near-twins, & DeepSeek R1 & DeepSeek V4 Flash are strangers. Worldview is invisible in code generation. In business analysis, forecasts, hiring, & policy work, it is a live input.
The Economist ran 25 frontier AI models through the World Values Survey1, the questionnaire that has mapped the moral beliefs of 100 countries since 1981. For this 2x2, there are two axes : first, traditional (religious) to secular. Second, survival, with a focus on collective basic needs, to self-expression & individualism.
Most models sit in the self-expression half of the map, which makes sense given the training data.
Surprisingly, the models are far apart. Gemini 3.1 Flash Lite & Qwen 3.6 Flash sit as neighbors, furthest in self-expression.
GPT-4o & DeepSeek R1 are near-twins, one trained in San Francisco, one in Hangzhou.
DeepSeek R1 & DeepSeek V4 Flash come from the same lab but lie at opposite ends of the secular / traditional axis.
Shared training data & similar labelers explain the near-twins. Different post-training choices explain the strangers. Common Crawl is 46% English2, so the base voice a model imitates is a college-educated American online. Anthropic then aligns Claude to principles from the UN Declaration of Human Rights3, a liberal document by construction.
Grok is off on its own, a traditional independent.
This variance changes the shopping list. Every RFP for an enterprise model today scores price, latency, context window, & benchmark scores. Worldview is not on the list. Should it be?
For code generation, SQL, log parsing, & image classification, that is fine. A computer program has no politics.
The moment a model is used for business decisions in a specific market, its worldview is a live input. Marketing copy, predictions of user behavior, & customer support tone all have to match the values of the target demographic.
AI worldviews have never been considered as part of AI procurement, but for certain use cases, it may need to become a consideration.