测量与缓解AI写作助手导致的作者形象扭曲

2026-04-27 12:00·55天前·Paul R\"ottger, Kobi Hackenburg, Hannah Rose Kirk, Christopher Summerfield

精选理由

这篇论文用近3000名作者和上万名读者的实验，量化了AI写作助手如何悄悄改变别人眼中的你。它揭示了一个所有用AI写东西的人都该知道的隐藏成本：工具在帮你表达的同时，也在重塑你的‘人设’。

AI 摘要

本研究通过三项大规模实验（2,939名作者、11,091名读者）评估AI写作助手对作者形象的影响。作者在有无AI协助下撰写政治观点段落，读者从29个社会感知维度进行盲评。结果显示，AI协助导致作者形象在所有维度发生扭曲：作者显得更固执己见、更有能力、情绪更积极，且其感知人口特征向特权群体偏移。尽管作者反对多数扭曲现象，却仍倾向于使用AI辅助文本。研究通过训练奖励模型在模型层面部分缓解了扭曲，但降低了用户接受度，表明AI写作助手的理想与非理想特性相互交织。这些扭曲在人类监督下依然普遍存在，可能对公共话语、信任与民主审议产生深远影响。

原文 · 未翻译

Computer Science > Computation and Language

[Submitted on 24 Apr 2026]

Title:Measuring and Mitigating Persona Distortions from AI Writing Assistance

Authors:Paul Röttger, Kobi Hackenburg, Hannah Rose Kirk, Christopher Summerfield

View PDF HTML (experimental)

Abstract:Hundreds of millions of people use artificial intelligence (AI) for writing assistance. Here, we evaluated how AI writing assistance distorts writer personas - their perceived beliefs, personality, and identity. In three large-scale experiments, writers (N=2,939) wrote political opinion paragraphs with and without AI assistance. Separate groups of readers (N=11,091) blindly evaluated these paragraphs across 29 socially salient dimensions of reader perception, spanning political opinion, writing quality, writer personality, emotions, and demographics. AI writing assistance produced persona distortions across all dimensions: with AI, writers seemed more opinionated, competent, and positive, and their perceived demographic profile shifted towards more privileged groups. Writers objected to many of the observed distortions, yet continued to prefer AI-assisted text even when made aware of them. We successfully mitigated objectionable persona distortions at the model level by training reward models on our experimental data (10,008 paragraphs, 2,903,596 ratings) to steer AI outputs towards faithful representation of writer stance. However, this came at a cost to user acceptance, suggesting an entanglement between desirable and undesirable properties of AI writing assistance that may be difficult to resolve. Together, our findings demonstrate that persona distortions from AI writing assistance are pervasive and persistent even under realistic conditions of human oversight, which carries implications for public discourse, trust, and democratic deliberation that scale with AI adoption.

Comments:
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2604.22503 [cs.CL]
	(or arXiv:2604.22503v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2604.22503 arXiv-issued DOI via DataCite

Submission history

From: Paul Röttger [view email]
[v1] Fri, 24 Apr 2026 12:31:11 UTC (1,354 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2026-04

Change to browse by:

References & Citations

Bookmark

Bibliographic and Citation Tools

Bibliographic Explorer (What is the Explorer?)

Connected Papers (What is Connected Papers?)

Litmaps (What is Litmaps?)

scite Smart Citations (What are Smart Citations?)

Code, Data and Media Associated with this Article

alphaXiv (What is alphaXiv?)

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub (What is DagsHub?)

Gotit.pub (What is GotitPub?)

Hugging Face (What is Huggingface?)

ScienceCast (What is ScienceCast?)

Demos

Replicate (What is Replicate?)

Hugging Face Spaces (What is Spaces?)

TXYZ.AI (What is TXYZ.AI?)

Recommenders and Search Tools

Influence Flower (What are Influence Flowers?)

CORE Recommender (What is CORE?)

Author
Venue
Institution
Topic

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

现象/趋势论文/研究

arXiv：cs.CL（计算语言学，全量分类）

精选69