# 清华等提出"表征平等"框架：评估大模型跨国模拟是否均衡

- 来源：OpenBMB (@OpenBMB)
- 发布时间：2026-08-24 22:00
- AIHOT 分数：51
- AIHOT 链接：https://aihot.virxact.com/items/cmt7b94zi20zuro73zdr1gm44
- 原文链接：https://x.com/OpenBMB/status/2091888261119828176

## AI 摘要

清华等高校提出“表征平等”框架，评估大语言模型作为跨国调查合成受访者时，模拟精度是否在各群体间均衡分布。测试覆盖59国、2420个亚组及9个开源模型，发现模拟不平等是系统性的：模型对澳大利亚、北美、欧洲模拟更准，且国家精度与人均GDP、互联网普及率正相关。

## 正文

LLMs are increasingly used as synthetic respondents in cross-country surveys. But do they simulate every country equally well, or just the rich, well-connected ones?

New in Computational Linguistics, from Tsinghua University with Fudan University, Nankai University, and Shanghai Jiao Tong University. It introduces Representational Equality: a framework for evaluating whether simulation accuracy is comparably distributed across populations, tested across 59 countries and 2,420 comparable subgroup units on 9 open models.

1️⃣ Representational equality is measurable, and it's not the same as accuracy. Simulation accuracy is 1−JSD relative to empirical World Values Survey distributions; representational equality is the cross-country spread, EqCV = σ/µ. GLM-4-9B is the most accurate (0.739) yet far from the most even (EqCV 0.0486), while ChatGLM3-6B leads on both equality (0.0286) and the joint AE score (0.096). Higher average accuracy does not imply higher representational equality.

2️⃣ The inequality is systematic and structural. Models simulate Australia, North America and Europe more accurately than the Middle East (Egypt, Jordan, Iraq). Country accuracy correlates positively with GDP per capita, Internet use, the Global Innovation Index and governance quality, and culturally with lower power distance and higher individualism and indulgence. Representational inequality thus follows systematic cross-country patterns.

3️⃣ One failure mode: asymmetric convergence toward a better-simulated country. On "how worried are you about losing your job or not finding a job," ~13% of Americans but over 75% of Mexicans answer "very much", yet Llama-3-8B predicts ~3% for both, much closer to the US pattern. Across about 2/3 questions its simulated Mexican distribution sits closer to the empirical US one than to Mexico’s.

4️⃣ What helps, and what doesn't. Additional-information supplementation most consistently improves both accuracy and representational equality (GLM-4-9B EqCV 0.0486→0.0255 with BM25 retrieval). Native-language prompting can improve many target groups, but the gains are uneven across languages and models. Continued post-training improves target-language averages unevenly, while preference alignment shows no systematic impacts; human-annotated data better preserve simulation accuracy than AI-annotated data.

Paper: https://doi.org/10.1162/COLI.a.648

#AI #THUNLP #OpenBMB #LLM #NLP #SocialSimulation #AIFairness #ComputationalLinguistics
