Anthropic 在一项版权和解协议中须向图书作者支付 15 亿美元。旧金山一家联邦法院批准了该和解协议,此前 Anthropic 在 2021 年至 2022 年间从盗版数据库 LibGen 和 PiLiMi 下载了图书。在大约 482,460 部列出的作品中,91.3% 被主张权利,每部作品约获赔 3,000 美元,是法定最低赔偿额的四倍。Anthropic 必须销毁这些盗版文件。作者保留对复现原作品的 AI 输出内容以及 Anthropic 未来行为的主张权利。这是集体诉讼历史上最大的一笔版权和解。
但这笔赔偿针对的是盗版行为,而非 AI 训练本身。法官 Alsup 此前曾裁定,在合法获取的图书上训练 AI 属于“变革性使用——而且是极其显著的变革性使用”,符合合理使用原则。未经作者同意大规模抓取互联网内容是否算作合法获取,仍是一个悬而未决的问题,因此合理使用的争论很可能远未结束。尽管如此,这项裁决对于在未经网站所有者同意的情况下抓取网络内容进行训练的 AI 实验室来说,看起来仍是一个里程碑,因为这类内容正是它们主要的训练数据来源。
法律 Justia
Anthropic has to pay book authors $1.5 billion in a copyright settlement. A federal court in San Francisco approved the settlement after Anthropic downloaded books from the piracy databases LibGen and PiLiMi between 2021 and 2022. Of roughly 482,460 listed works, 91.3 percent were claimed, netting about $3,000 each, four times the statutory minimum. Anthropic must destroy the pirated files. Authors retain claims over AI outputs that reproduce original works and over Anthropic's future conduct. It's the largest copyright settlement in class action history.
But the payout covers piracy, not AI training itself. Judge Alsup had previously ruled that training AI on legally obtained books is "transformative - spectacularly so" and falls under fair use. Whether mass scraping of internet content without authors' consent counts as legal acquisition remains an open question, so the fair use debate is likely far from over. Still, the ruling looks like a milestone for AI labs that trained on web content without website owners' consent, their main source of training data.
Law Justia