核心要点
- 包括索尼音乐和华纳音乐在内的多家大型音乐出版商已对Anthropic及其高管个人提起诉讼。
- 这家AI公司被指控未经许可使用数万首受版权保护的音乐作品(主要是歌词)来训练其Claude模型。
- 本案的核心在于Anthropic如何获取训练数据,而不仅仅是它如何使用这些数据。
大型音乐出版商指控Anthropic非法下载并使用数万首受版权保护的音乐作品来训练Claude。
索尼音乐、华纳音乐等出版商已在加州北区联邦法院对Anthropic提起诉讼。根据起诉书,CEO达里奥·阿莫代伊和联合创始人本杰明·曼被列为个人被告,理由是他们涉嫌指挥和监督受版权保护文件的种子下载行为。
这份48页的起诉书写道:“阿莫代伊博士明确指示、批准、控制并故意诱导曼先生及其他Anthropic员工实施这些侵权行为。”原告称这是“历史上规模最大、最明目张胆的知识产权持续盗窃行为之一”。
起诉书聚焦于音乐作品(主要是歌词),以及乐谱和衍生作品。原告要求对每件被侵权作品最高赔偿15万美元,并对非法删除版权管理信息(包括版权声明和其他识别信息)的每次违规行为最高赔偿2.5万美元。
Anthropic此前已在这场斗争中败诉过一次
2025年9月,Anthropic同意支付15亿美元,与美国历史上最大规模的版权和解——向作者和出版商赔偿其在AI训练中使用盗版书籍的费用。在那起案件中,让Anthropic败诉的并非使用受版权保护的数据进行训练,而是通过非法种子下载获取这些数据。
索尼-华纳诉讼瞄准的正是同一个软肋。Anthropic被指控从盗版图书馆LibGen和PiLiMi种子下载了至少700万本书籍。该诉讼将这种下载行为本身视为独立的版权侵权行为,无论这些作品是否最终进入了商业化的Claude模型。
该公司还被指控未经出版商同意,从 MusixMatch 和 LyricFind 等授权平台抓取歌词,违反了这些平台的服务条款。该诉讼对 Anthropic 使用 Books3、The Pile 和 Common Crawl 等数据集的行为提出质疑,这些数据集据称包含未经授权的内容。诉讼还指控该公司扫描并销毁了二手歌本和乐谱收藏。
合成数据作为版权漏洞
原告还针对 Anthropic 在训练过程中如何处理盗版内容提出指控。Anthropic 已公开否认直接使用 LibGen 和 PiLiMi 的书籍来训练商业版 Claude 模型,但诉状认为,这一否认取决于 Anthropic 如何定义“训练”。
原告指控 Anthropic 至少训练了一个商业版 Claude 模型,其训练数据是由一个非商业模型生成的合成数据,而该非商业模型本身是从 LibGen 和 PiLiMi 的文本中学习而来的。他们还声称 Anthropic 使用这样一个模型为商业版 Claude 模型提供强化反馈。这些指控的完整范围将在证据开示阶段揭晓。
在另一起相关案件中,慕尼黑地区法院于 2025 年 11 月裁定,受版权保护的歌词即使存储在模型参数中,也构成复制行为,而聊天机器人输出这些歌词则构成非法公开传播。法院认定模型运营方应对其模型生成的内容负责,而非输入提示词的用户,即使这些提示词是专门为生成受版权保护的歌词而设计的。
Courtlistener
Key Points
- Major music publishers including Sony Music and Warner Music have sued Anthropic and its executives personally.
- The AI company allegedly used tens of thousands of copyrighted compositions, mostly song lyrics, to train its Claude models without permission.
- The case centers on how Anthropic acquired the training data, not just how it used it.
Major music publishers accuse Anthropic of illegally downloading and using tens of thousands of copyrighted musical compositions to train Claude.
Sony Music, Warner Music, and other publishers have sued Anthropic in federal court in Northern California. According to the complaint, CEO Dario Amodei and co-founder Benjamin Mann are named as individual defendants for their alleged role in directing and overseeing the torrenting of copyrighted files.
"Dr. Amodei expressly directed, approved, controlled, and intentionally induced these infringements by Mr. Mann and other Anthropic employees," the 48-page complaint states. The plaintiffs call it "one of the largest and most blatant ongoing thefts of intellectual property in history."
The complaint focuses on musical compositions, mostly song lyrics, plus sheet music and derivative works. The plaintiffs are seeking up to $150,000 per infringed work and up to $25,000 per violation for unlawful removal of copyright management information, including copyright notices and other identifying information.
Anthropic has already lost this fight once
In September 2025, Anthropic agreed to the largest copyright settlement in U.S. history, paying $1.5 billion to authors and publishers for using pirated books during AI training. What sank Anthropic in that case wasn't using copyrighted data for training, but acquiring it through illegal torrent downloads.
The Sony-Warner lawsuit targets the same weak spot. Anthropic allegedly torrented at least seven million books from the pirate libraries LibGen and PiLiMi. The lawsuit treats that downloading as a standalone copyright infringement, regardless of whether the works ever made it into a commercial Claude model.
The company also allegedly scraped song lyrics from licensed platforms like MusixMatch and LyricFind without publisher consent, violating those platforms' terms of service. The complaint challenges Anthropic's use of datasets like Books3, The Pile, and Common Crawl, which allegedly contain unauthorized content. It also accuses the company of scanning and destroying used songbooks and sheet music collections.
Synthetic data as a copyright loophole
The plaintiffs also go after how Anthropic handled torrented content during training. Anthropic has publicly denied using LibGen and PiLiMi books directly to train commercial Claude models, but the complaint argues that denial depends on how Anthropic defines "training."
The plaintiffs allege Anthropic trained at least one commercial Claude model on synthetic data generated by a non-commercial model that had itself learned from LibGen and PiLiMi texts. They also claim Anthropic used such a model to provide reinforcement feedback to a commercial Claude model. The full scope of these claims will come out during discovery.
In a related case, the Munich Regional Court ruled in November 2025 that copyright-protected song lyrics count as reproductions even within a model's parameters and that chatbot output of those lyrics amounts to unlawful public disclosure. The court held model operators responsible for what their models produce, not the users who prompted the output, even when those prompts were specifically designed to generate copyrighted lyrics.
Courtlistener