一群出版商和作者已对谷歌提起集体诉讼,指控这家科技巨头使用其受版权保护的作品来训练其 AI 平台 Gemini。
原告群体包括 Hachette、Cengage、Elsevier、作家 Scott Turow 以及 S.C.R.I.B.E.。根据诉讼文件,他们还指控谷歌故意删除或更改这些作品上的版权信息,以“掩盖……其 Gemini 模型是在被盗材料上训练的”这一事实。
这起诉讼只是出版商、作者及其他版权持有者针对谷歌、Meta、OpenAI 和 Anthropic 等 AI 公司提起的众多投诉之一。
尽管其中许多诉讼仍在审理中,但加利福尼亚州法院的两项早期裁决对 AI 公司有利,裁定在互联网出现之前就已存在且此后未再更新的美国版权法下,使用受版权保护的作品进行 AI 训练属于“合理使用”。
然而,Anthropic 因盗用其训练所用的作品而被处以 15 亿美元罚款,这是美国版权法历史上金额最高的赔偿。大约有 50 万名作家有资格获得至少 3000 美元的赔偿。不过,许多作者选择不接受和解,以便就 AI 训练问题采取进一步的法律行动。
加州法官的裁决对于其他法院如何看待科技公司的合理使用辩护来说并非好兆头,但这场冲突过于微妙,这些裁决尚不足以确立无可争议的先例。针对谷歌的诉讼是在美国纽约南区联邦地区法院提起的,这给了另一位法官发表意见的机会。
在谷歌的案例中,出版商与该公司之间存在着更为微妙且长期的合作关系。诉讼指出,出版商和作者长期以来一直向谷歌提供受版权保护的作品,其特定目的是通过谷歌图书(Google Books)使这些书籍可供搜索。这些搜索结果不允许用户查看整本书,而是提供书籍的简短片段以及书目信息。原告声称,谷歌使用这些书籍的副本以及上传至谷歌应用商店(Google Play store)的书籍来训练 Gemini,尽管它从未获得这样做的许可。
诉讼文件写道:“谷歌明知缺乏授权,却非法复制了所有这些限定范围项目中的作品用于人工智能训练。”
原告还引用了谷歌的一份内部文件,该文件据称指出,将受版权保护的书籍用于人工智能训练可能“对谷歌来说问题极其严重”,并可能导致“100亿至1000亿美元的潜在罚款”。
谷歌未立即回应置评请求。
A group of publishers and authors have filed a class action lawsuit against Google, accusing the tech giant of using their copyrighted works to train its AI platform, Gemini.
The group of plaintiffs, which includes Hachette, Cengage, Elsevier, author Scott Turow, and S.C.R.I.B.E., also alleges that Google intentionally removed or changed copyright information on these works to “conceal… that its Gemini Models were trained on stolen materials,” according to the lawsuit.
This lawsuit is just one of many complaints that publishers, authors, and other copyright holders have filed against AI companies such as Google, Meta, OpenAI, and Anthropic.
While many of these lawsuits are still pending, two early court decisions in California have favored the AI companies, ruling that the use of copyrighted works for AI training is considered “fair use” under U.S. copyright law that has not been updated since before the existence of the internet.
Anthropic was, however, fined $1.5 billion for pirating the works it trained on, marking the largest payout in the history of U.S. copyright law. Around half a million writers were eligible for payments of at least $3,000. However, many authors opted out of receiving the settlement so that they could pursue further legal action over AI training.
The California judges’ decisions don’t bode well for how other courts may view the tech companies’ fair use defense, but the conflict is too nuanced for these rulings to establish an inarguable precedent. The lawsuit against Google was filed in the U.S. District Court for the Southern District of New York, giving a different judge the opportunity to weigh in.
In the Google case, the publishers have a more nuanced, long-term relationship with the company. The lawsuit explains that publishers and authors have a long history of providing Google with copyrighted works for the specific purpose of making books searchable through Google Books. These search results do not allow users to view entire books. Instead, they provide access to short snippets of the book along with bibliographic information. The plaintiffs claim that Google trained Gemini on copies of these books, as well as books uploaded to the Google Play store, even though it never received permission to do so.
“Google illegally copied works from all these scope-limited programs for AI training, knowing it lacked authorization to do so,” the lawsuit reads.
The plaintiffs also cite an internal document from Google that allegedly states that using copyrighted books for AI training could be “highly problematic for Google” and might result in “$10Bs-$100Bs in potential fines.”
Google did not immediately respond to a request for comment.