全球几家领先的音乐出版商认为,Anthropic 在一项历史性和解中受到的惩罚太轻了——在这项和解中,这家 Claude 开发商在承认盗用超过 700 万本书籍用于训练 AI 后,向作者支付了 15 亿美元。
音乐出版商在上周五提起的诉讼中表示:“15 亿美元显然不足以震慑一家公司的侵权行为,这家公司已把这种大规模侵权转化为高达 2 万亿美元的惊人估值。”
起诉 Anthropic 的音乐出版商包括索尼、EMI 和华纳查普尔。他们指控 Anthropic 的非法盗版行为还包括“成千上万”本含有歌词和乐谱的书籍,这些书籍涉及他们数百首甚至更多受版权保护的音乐作品。
诉状称,越来越多在音乐排行榜上与 AI 生成作品竞争的词曲作者正受到这种盗版行为的伤害。
Anthropic 据称计划“永久”保留这些盗版书籍用于训练 AI,其中包括收录了披头士乐队全部作品、泰勒·斯威夫特“最佳”歌曲以及“VH1 摇滚乐百大金曲”的书籍。出版商指控称,如果没有禁令来终止不当的 AI 训练,Anthropic 将继续侵犯音乐版权——同时通过逐字复制歌词、生成模仿其最受欢迎作品“核心”的“新”歌曲,不正当地抢占艺术家在市场中的位置。例如,Anthropic 可能会依赖歌本或乐谱来响应用户的提示词,比如要求 Claude 把埃米纳姆的歌词改写成一首“碧昂丝风格”的歌曲。
音乐出版商指控称:“这是一场肆无忌惮的大规模非法盗版、抓取和下载受版权保护作品的活动,目的是开发、运营并攫取巨额利润。”
索尼重新翻出 Anthropic 员工的聊天记录。
Anthropic 的“大规模非法种子下载活动”始于 2021 年 7 月,当时 Anthropic 联合创始人 Benjamin Mann“亲自使用 BitTorrent 通过种子下载非法下载并上传了数百万本盗版书籍”,这些书籍来自 Library Genesis(LibGen)——一个备受争议的盗版图书馆。Anthropic 联合创始人兼 CEO Dario Amodei 据称批准了此次种子下载,他和 Mann 均被音乐出版商诉讼列为个人被告。
不过,投诉指出,到 2021 年底,FBI 已关闭了 LibGen,“但在网络盗版者复制其内容创建新图书馆之前并未关闭”,该新图书馆名为“Z-Library”。诉讼称,尽管 Z-Library 也很快被关闭,但图书作者诉讼中披露的内部消息显示,Anthropic 通过“Pirate Library Mirror”(PiLiMi)获取了这份副本的副本。
Mann 被指控在 PiLiMi 上线后不久即指示员工进行种子下载,并向同事发送消息称该镜像“来得正是时候!”作为回应,一名 Anthropic 员工回复道:“zlibrary 我的挚爱。”
音乐出版商指出这些评论是 Anthropic 宣扬使用盗版并依赖种子下载快速获取新训练数据的证据。他们声称,总计而言,爬取 LibGen 和 PiLiMi 的非机密目录显示,“书目元数据如标题、作者和 ISBN”表明“Anthropic 种子下载了至少数百本包含乐谱和歌曲歌词的书籍”,这些作品归起诉的出版商所有。
投诉称,通过证据开示程序,出版商声称将揭示 Anthropic 种子下载行为的“全部范围”。
出版商怀疑 Claude 在热门歌曲上进行了训练
出版商承认,Anthropic 否认使用其从 LibGen 和 PiLiMi 种子下载的任何书籍来训练商业 Claude AI 模型。然而,他们声称,Anthropic 的声明取决于如何定义“训练”,他们认为如果法院追溯 Anthropic 的步骤足够远,可能会发现商业模型在某个时间点确实在盗版作品上进行了训练。正如他们所解释的:
“通常,AI 开发包含一个‘预训练’阶段,其中一个 AI 模型可能会使用另一个 AI 模型生成的‘合成数据’进行训练,或接收来自另一个 AI 模型的其他行为反馈。Anthropic 已使用由非商业 AI 模型创建的合成数据训练了至少一个其商业发布的 Claude 模型,该非商业模型基于源自 LibGen 和/或 PiLiMi 的文本进行训练。Anthropic 至少使用了一个基于源自 LibGen 和/或 PiLiMi 的文本训练的非商业模型,为至少一个商业 Claude 模型提供强化反馈。”
此外,他们的诉状指出,近期在图书作者案件中解封的文件“显示 Anthropic 利用其从 LibGen 非法下载的盗版书籍来辅助其安全护栏”。例如,在一份文件中,Anthropic 的一位证人作证称,Anthropic 已停止在 LibGen 上训练大语言模型(LLM),但继续使用 LibGen 数据集来检查输出中的长文本串是否与源文本过于接近。
Anthropic 的一位发言人向 Ars 提供了一份声明,该声明似乎暗示音乐出版商是在通过翻查图书作者已和解的诉状来抓救命稻草。
“这是同一批律师提起的第三起诉讼,不过是翻炒已在法院审理的案件中的指控,”Anthropic 的发言人表示。“训练生成式 AI 模型是变革性的合理使用——正如法院在 Bartz 案中所裁定的——我们将坚决为自己辩护。”
据称受 AI 歌曲伤害的词曲作者
Anthropic 的声明没有提到,合理使用的裁定取决于图书作者无法证明市场损害,也无法证明 AI 工具已在市场上取代了他们。
诉讼称,只有 Anthropic 自己最清楚用户有多频繁地依赖 Claude 来生成音乐人歌曲的替代品。但出版商似乎认为,音乐版权方在证明市场损害方面可能比图书作者更有胜算。
他们声称,Anthropic 公开追踪其AI产品对歌词作者构成的威胁,并且该公司知道,目前登顶音乐排行榜的AI生成歌曲,正在与那些作品未经付费即被用于训练AI、从而被挤下榜首位置的音乐人直接竞争。诉讼称,最引人注目的是,美国版权局“观察到,‘当一个生成式AI模型的输出——即使与某一特定受版权保护的作品不构成实质性相似——在该类作品的市场中与之竞争时’(包括歌词的情况),这些输出可能会稀释版税池”。
具体而言,审理图书作者案件的法官写道:“就像任何渴望成为作家的读者一样,Anthropic的大语言模型在受版权保护的作品上进行训练,不是为了抢跑、复制或取代它们——而是为了转过一个艰难的弯,创造出不同的东西。”但法院是否会认同这一逻辑同样适用于词曲作者或代表他们的音乐出版商,目前尚不清楚。
音乐出版商指控称,Claude被有意训练成会复述歌词,其依赖的不仅是盗版文库,还有其他未经授权的数据集,包括在未获权利人许可的情况下从歌词网站抓取的盗版内容。此外,Anthropic还销毁实体书籍以获取更多数据,据称制作了“数百本歌本和乐谱集”的未经授权数字副本,这也构成了版权侵权。
据出版商称,Claude模型在被提示时会调取歌词,且不附带版权管理信息,甚至在用户并未要求时也会生成歌词。例如,如果用户向Claude询问某首特定歌曲的和弦进行,“该AI模型往往会生成包含‘与这些和弦并列的受版权保护歌词’的输出”,诉状称。诉讼还指出,Claude在制作“新”歌曲时,也常常将真实歌词与AI生成的歌词混搭在一起,并且它能够“针对各种各样的用户提示,再现这些作品的‘核心’”。
Anthropic 似乎是刻意让模型以这种方式表现的。据称,在 Anthropic 要求员工测试 AI 模型能否根据某人喜爱的音乐推荐风格相近的歌曲之后,Anthropic 的员工“反复向 AI 模型提示音乐出版商拥有版权的歌词,鼓励模型生成包含这些歌词的输出内容”。
Amodei 在图书作者诉讼中作证时表示,Anthropic 本可以合法购买受版权保护的作品,但最终却选择了盗版下载,原因是 Anthropic 想避免一场“法律/实践/商业上的拉锯战”。法院将 Anthropic 的意图概括为下载盗版书籍“以避免付费的麻烦”,而音乐出版商则强调,Anthropic 从未像其最大的 AI 竞争对手那样主动与它们洽谈授权协议。
Anthropic 对其训练数据来源严格保密,但音乐出版商认为法院应要求其提高透明度。出版商要求,应责令 Anthropic“提供其 AI 模型的训练数据、训练方法及已知能力的详细说明”。这样一来,各类媒体的艺术家都能评估 Anthropic 是如何获取其作品的、哪些作品被纳入训练,以及这些作品对商业模型可能产生了多大影响。
音乐出版商声称,如果放任不管,Anthropic 依赖盗版和未经授权的作品来训练 AI,会让如今词曲作者更难以此为生,同时还会用低质量的 AI 仿制版知名歌曲来稀释市场。当然,诉状还指出,盗版也让出版商更难通过授权歌曲供 AI 训练来获得报酬。
“Anthropic 从未寻求也未获得任何许可,以合法利用音乐出版商的受版权保护作品用于任何用途,更不用说用于 AI 训练数据或输出内容了,”他们的诉状写道。“Anthropic 通过无偿利用音乐出版商及其词曲作者的劳动成果来为自己牟利,这削弱了他们投资、支持并拓展当前及未来创作事业的动力。”
![]()
![]()
Some of the world’s leading music publishers think that Anthropic got off too light in a historic settlement where the Claude maker paid authors $1.5 billion after admitting to pirating more than 7 million books to train AI.
“$1.5 billion is obviously not a large enough settlement to deter infringing conduct by a company that has parlayed such mass infringement into a staggering $2-trillion-dollar valuation,” music publishers said in a lawsuit filed Friday.
Music publishers suing Anthropic include Sony, EMI, and Warner Chappell. They alleged that Anthropic’s illegal torrenting also included “thousands upon thousands” of books containing lyrics and sheet music to hundreds or more of their copyrighted musical compositions.
Songwriters who are increasingly competing with AI-generated works in music charts are harmed by that piracy, their complaint said.
Pirated books that Anthropic allegedly plans to keep “forever” to train AI include titles featuring the complete works of the Beatles, Taylor Swift’s “best” songs, and “VH1’s 100 Greatest Songs of Rock & Roll.” Without an injunction ending the improper AI training, Anthropic will continue to violate music copyrights—while improperly substituting artists in their markets by reproducing song lyrics verbatim and by generating “new” songs that mimic the “heart” of their most popular works, publishers alleged. For example, Anthropic may rely on songbooks or sheet music to respond to prompts asking Claude to change Eminem lyrics into a song written in “Beyonce’s style.”
“It’s a brazen campaign of illegally torrenting, scraping, and downloading copyrighted works on a massive scale in order to develop, operate, and reap enormous profits,” music publishers alleged.
Sony resurfaces Anthropic staff chats
Anthropic’s “mass campaign of illegal torrenting” began in July 2021, when Anthropic co-founder Benjamin Mann “personally used BitTorrent to unlawfully download and upload via torrenting millions of pirated books” from Library Genesis (LibGen), a controversial pirate library. Dario Amodei, Anthropic co-founder and CEO, allegedly approved the torrenting, and both he and Mann are named individually as defendants in music publishers’ lawsuit.
By the end of 2021, though, the FBI had shut down LibGen, the complaint noted, “but not before online pirates copied its contents to create a new library,” named “Z-Library.” And although Z-Library also got shut down quickly, internal messages revealed in the book authors’ fight showed that Anthropic got access to a copy of the copy through the “Pirate Library Mirror” (PiLiMi), the lawsuit said.
Mann is accused of directing employees to torrent PiLiMi soon after it became available, messaging colleagues that the mirror dropped “just in time!” Responding, an Anthropic staffer wrote back, “zlibrary my beloved.”
Music publishers pointed to these comments as proof that Anthropic extolled the use of piracy and relied on torrenting to quickly access new training data. In total, they alleged that crawling through LibGen and PiLiMi’s non-confidential catalogs showed “bibliographic metadata like title, author, and ISBN” that indicated that “Anthropic torrented at least hundreds of books containing sheet music and song lyrics to musical compositions” owned by publishers suing.
Through discovery, publishers claim they will reveal the “full extent” of Anthropic’s torrenting, the complaint said.
Publishers suspect Claude trained on hit songs
Publishers acknowledged that Anthropic denies using any of the books that they torrented from LibGen and PiLiMi to train commercial Claude AI models. However, they claimed that Anthropic’s statements depend on how “training” is defined, and they think that if the court traces Anthropic’s steps back far enough, it may reveal that commercial models were at some point trained on pirated works. As they explained:
“Often, AI development includes a ‘pretraining’ phase where one AI model may train on ‘synthetic data’ created by another AI model or receive other behavioral feedback from another AI model. Anthropic has trained at least one of its commercially released Claude models using synthetic data created by a non-commercial AI model that was trained on text derived from LibGen and/or PiLiMi. Anthropic employed at least one non-commercial model that was trained on text derived from LibGen and/or PiLiMi to provide at least one commercial Claude model with reinforced feedback.”
Additionally, their complaint noted that recently unsealed documents from the book authors’ case “revealed that Anthropic exploits the pirated books it illegally torrented from LibGen in connection with its guardrails.” For example, in one filing, an Anthropic witness testified that Anthropic stopped training large language Models (LLMs) on LibGen but continued using the LibGen dataset to see if long strings of text in outputs too closely matched source text.
A spokesperson for Anthropic provided Ars with a statement that seems to suggest that music publishers are grasping at straws by digging through book authors’ settled complaint.
“This is the third lawsuit from the same lawyers, recycling allegations from cases already before the courts,” Anthropic’s spokesperson said. “Training generative AI models is a transformative fair use—as the court held in Bartz—and we will defend ourselves robustly.”
Songwriters allegedly harmed by AI songs
Anthropic’s statement neglects to mention that the fair use ruling hinged on book authors’ inability to prove market harms or that AI tools had substituted them in their markets.
Only Anthropic knows for sure how often users rely on Claude to generate substitutes for musicians’ songs, the lawsuit said. But publishers seem to think that music rightsholders may have a better chance at proving market harms than authors did.
They’ve alleged that Anthropic publicly tracks the threat to song lyricists from its AI products and that the company knows that AI-generated songs currently topping music charts compete directly with musicians whose songs were used without payment to allegedly train the AI replacing them in top slots. Most glaringly, the US Copyright Office has “observed that ‘where a generative AI model’s outputs, even if not substantially similar to a specific copyrighted work, compete in the market for that type of work,’ including in the case of song lyrics,” the outputs may dilute royalty pools, the lawsuit said.
Specifically, the judge in the book authors’ case wrote that, “like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them—but to turn a hard corner and create something different.” But it’s unclear if courts will agree that the same holds true for songwriters or the music publishers who represent them.
Music publishers alleged that Claude was intentionally trained to regurgitate lyrics, relying not just on pirate libraries but also on other unauthorized datasets, including pirated content scraped from lyrics sites without the permission of rights holders. Additionally, Anthropic destroyed physical books to harvest more data, allegedly creating unauthorized digital copies “of hundreds of songbooks and sheet music collections,” which also violated copyrights.
According to publishers, Claude models will fetch lyrics when prompted without sharing copyright management information, and they even generate lyrics when users do not request them. For example, if a user asks Claude for a particular song’s chord progression, “the AI model will often generate output” containing “copyrighted lyrics alongside those chords,” the complaint said. And Claude will also often mash up actual lyrics with AI-generated lyrics when making “new” songs, the lawsuit said, and it’s capable of “reproducing the ‘heart’ of those works in response to a wide range of user prompts.”
Anthropic seemingly designed the model to perform this way. Allegedly, Anthropic workers “repeatedly prompted the AI models for Music Publishers’ copyrighted lyrics, encouraging the models to generate output containing those lyrics” after Anthropic asked them to test if the AI models could recommend comparable songs based on someone’s favorite music.
Amodei testified during the book authors’ litigation that Anthropic could have legally purchased copyrighted works but torrented them instead, because Anthropic wanted to avoid a “legal/practice/business slog.” The court summarized Anthropic’s intent as downloading pirated books “to avoid the trouble of paying for them,” and music publishers emphasized that Anthropic never approached them to strike licensing deals as its biggest AI rivals have.
Anthropic closely guards its training data sources, but music publishers think the court should require more transparency. Anthropic should be ordered to “provide an accounting of the training data, training methods, and known capabilities of Anthropic’s AI models,” publishers demanded. Then artists in all types of media could assess how Anthropic acquired their works, what works were ingested, and how much they may have influenced a commercial model.
Left unchecked, music publishers alleged that Anthropic’s reliance on pirated and unauthorized works to train AI makes it harder to make a living as a songwriter today, while diluting the market with low-quality AI copies of recognizable songs. And of course, the complaint noted that piracy also makes it harder for publishers to get paid to license songs to train AI.
“Anthropic has never sought nor obtained any license to lawfully exploit Music Publishers’ copyrighted works for any use, let alone for AI training data or output,” their complaint said. “Anthropic enriches itself through the uncompensated exploitation of Music Publishers’ and their songwriters’ labor, reducing their incentive to invest in, support, and expand present and future creative efforts.”
![]()
![]()
