Rohan Paul@rohanpaul_ai
36AI 编辑部评分,满分 100

LLM 危害分类研究:五阶段风险全景图

2026-08-13 21:30· 37分钟前
AI 导读

一项基于约 200 篇论文与事件报告的研究,构建了大语言模型危害的完整分类图谱,按模型生命周期划分为 5 个阶段。研究覆盖发布前的数据隐私与能源消耗、输出中的偏见与幻觉、恶意滥用、社会层面影响及行业应用风险,并梳理了现有技术与政策防御手段,强调多层防护的必要性。论文题为“LLM Harms: A Taxonomy and Discussion”,发布于 arXiv。

This study builds a simple map of all the main ways large language models can cause harm.

Using about 200 research papers and incident reports, it groups harms into 5 stages of a model's life.

Before the model is released, it highlights issues like scraped personal data, training on text without people's consent, heavy energy use, and low paid annotation work.

In what the model writes out, it focuses on biased or stereotyped language, toxic or false content, and hallucinations that look trustworthy.

For intentional misuse, it shows how people can generate scams, targeted abuse, propaganda, and even prompt based attacks on connected systems.

At the broader social level, it describes job disruption, political manipulation, concentration of computing power, and unequal access to advanced models across regions and languages.

When models are built into tools for healthcare, finance, education, or creative work, it explains how errors and bias can quietly shape real decisions.

Across all stages, the authors line up existing technical and policy defenses and argue that only many layered safeguards together can keep risks manageable.

  • arxiv. org/abs/2512.05929

Paper Title: "LLM Harms: A Taxonomy and Discussion"

来源:Rohan Paul · x.com

LLM 危害分类研究:五阶段风险全景图

Rohan Paul · @rohanpaul_ai · X·2026-08-13 21:30·37分钟前
AI 导读

一项基于约 200 篇论文与事件报告的研究,构建了大语言模型危害的完整分类图谱,按模型生命周期划分为 5 个阶段。研究覆盖发布前的数据隐私与能源消耗、输出中的偏见与幻觉、恶意滥用、社会层面影响及行业应用风险,并梳理了现有技术与政策防御手段,强调多层防护的必要性。论文题为“LLM Harms: A Taxonomy and Discussion”,发布于 arXiv。

This study builds a simple map of all the main ways large language models can cause harm.

Using about 200 research papers and incident reports, it groups harms into 5 stages of a model's life.

Before the model is released, it highlights issues like scraped personal data, training on text without people's consent, heavy energy use, and low paid annotation work.

In what the model writes out, it focuses on biased or stereotyped language, toxic or false content, and hallucinations that look trustworthy.

For intentional misuse, it shows how people can generate scams, targeted abuse, propaganda, and even prompt based attacks on connected systems.

At the broader social level, it describes job disruption, political manipulation, concentration of computing power, and unequal access to advanced models across regions and languages.

When models are built into tools for healthcare, finance, education, or creative work, it explains how errors and bias can quietly shape real decisions.

Across all stages, the authors line up existing technical and policy defenses and argue that only many layered safeguards together can keep risks manageable.

  • arxiv. org/abs/2512.05929

Paper Title: "LLM Harms: A Taxonomy and Discussion"

来源:Rohan Paul· x.com