# 隐藏 Airtag 追踪显示亚马逊正在销毁珍本图书以训练 AI

- 来源：Ars Technica：AI（RSS）
- 作者：Ashley Belanger
- 发布时间：2026-08-18 02:13
- AIHOT 分数：55
- AIHOT 链接：https://aihot.virxact.com/items/cmsxkc5e5035mroz0ev5qge2y
- 原文链接：https://arstechnica.com/tech-policy/2026/08/hidden-airtag-reveals-amazon-is-trashing-rare-books-to-train-ai

## AI 摘要

404 Media 调查发现，书商在珍本图书中植入 Airtag，追踪到亚马逊位于拉斯维加斯的 AI 训练设施 VGT3，该团队将书籍拆解并扫描页面。亚马逊拒绝就此事置评，仅称通过商业渠道购书以改进产品。调查还证实了书商理论：AI 公司按 ISBN 清单系统扫描图书，且亚马逊曾因图书供应短缺而担忧仓库关闭。

## 正文

For the past year or so, booksellers have suspected that AI firms are buying up huge lots of rare books, then destroying them after scanning them to train AI. But this was hard to prove until now, as 404 Media reports that an Airtag hidden in a rare book shows that at least one tech giant, in the race to advance its frontier models, is behind some of the bulk orders: Amazon.

On Monday, 404 Media revealed that it had connected with a bookseller who agreed to plant an Airtag in a rare book that was part of a bulk order. That Airtag was then tracked to an Amazon AI training facility in Las Vegas that housed a team focused on tearing books from their spines and scanning pages, 404 Media reported. Apparently tone-deaf to the escalating backlash over destructive book scanning, a logo on the door of that team’s warehouse, VGT3, showed a Tyrannosaurus rex preparing to devour a book, 404 Media documented.

Amazon deflects

Amazon declined to comment on 404 Media’s findings, only providing Ars with the same statement it gave to 404 Media, which does not mention AI training specifically.

“Amazon purchases books through commercial channels to help develop and improve the products and services our customers use,” Amazon’s statement said.

Ars Video

How Lighting Design In The Callisto Protocol Elevates The Horror

However, Amazon is developing what it considers frontier AI models, which require a massive amount of unique training data to stay competitive with leading firms like Google, OpenAI, or Anthropic. Right now, firms carefully guard their training data to avoid losing an edge. And training models on text from rare books that are difficult to find would seemingly offer an advantage for Amazon, especially since rivals like Anthropic and xAI have publicly stated that they are not training on rare or antique books.

Further, it seems that Amazon needed a new source of original text. 404 Media flagged discussions in online forums where VGT3 workers suggested that earlier this year Amazon had run low on books to scan. The shortage was so alarming that they worried the warehouse might shut down if Amazon couldn’t find more books to scan. At one point, the supply completely ran out, workers said. But the facility is still operational, 404 Media reported. And it’s now confirmed that bulk orders of rare books delivered there are systematically destroyed by this crew.

Many book lovers are horrified by destructive book-scanning (since there is an alternative), with one staunch critic, Michael Burry, even reportedly labeling the practice to be “evil incarnate.” However, Amazon workers reported in forums that they consider the gig to be a “nice” opportunity for those drawn to a humdrum job with flexible hours.

Debate rages over rare books

On top of revealing that rare works that booksellers value are getting chewed up and swallowed by Amazon’s AI machine, 404 Media suggested that its investigation helped firm up another bookseller theory about why AI firms might be ordering certain rare books and not others.

After a reportedly historic year of sales, booksellers had suspected that AI firms were targeting books with ISBN numbers in order to ensure that the highest volume of unique works were present in training data sets. And 404 Media’s review of Amazon workers’ online discussions indicated that they were trained to scan barcodes or ISBNs before scanning books. That practice, 404 Media reported, “gives further credence” to booksellers’ theory that “AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs.”

For booksellers, the money may be good, but the risk that their carefully sourced collections will be destined for destructive book scanning like Amazon’s raises an ethical dilemma. They know how to assess a wide range of rare books to determine their value, and AI firms seem to be skipping that step in hunting low-cost, unique ISBNs to complete their checklists.

Right now, the books that AI firms are apparently buying up aren’t necessarily the kind of prized first editions of celebrated works that are typically valued quite highly. Instead, AI firms often target older books with lower monetary value, such as books that were never translated from a foreign language that’s not widely used today or books that were never popular enough to be widely distributed.

However, these works may still have “historical value, intellectual value, sentimental value” that AI firms overlooked, the bookseller who planted the Airtag told 404 Media. A rare book’s value can be derived from “all sorts of things” that “the AI companies don’t care about. They just want the content as a bunch of words strung together.”

Redditors weigh in

On Reddit, some book fans debated whether it was that problematic that companies are destroying rare books to train AI, especially since, as the BBC reported, some of these books have been sitting on booksellers’ shelves for decades gathering dust.

“You were not going to buy that old paper book,” one Redditor commented in response to a post lamenting that “an obscure book from 1700 is now a museum piece and may reveal day-to-day stuff that we didn’t know.”

In that thread, the original poster said that the real problem was that tiny details and even major historical insights that can be gleaned from reviewing rare books will be lost to AI greed. AI firms will “never share the contents” of books they scan, the poster said, “as they don’t want anyone else to be able to train their AI” on the same works.

“This sounds like propaganda from the AI haters,” another Redditor pushed back, but a subsequent commenter shared similar fears. Although people might clash with someone who argues that training AI on these works will make all the knowledge that a work contains more accessible online, the commenter suggested instead that AI models would only make available “warped, censored, and paywalled fragments of ideas” from forgotten works.

“Even if you love the technology, you can admit that the concept of an AI literally eating books to become more powerful is pretty dystopian,” someone else on the thread said.

Some booksellers agree that not every text needs to be saved, the BBC reported. However, Scottish bookseller Derek Walker told the BBC that AI firms should be striving to distinguish between works that won’t be missed much—such as little-known academic texts that may technically be rare—and lesser-known antique works that may be “the only known surviving example of an edition.”

“It would be a much more significant problem if one like that were to be bought for destruction, having survived this long,” Walker said.

For AI firms guarding their training data and using services that mask their identities as buyers, there’s likely little desire to discuss their bulk buying directly with booksellers. Allowing booksellers to weigh in on the works they plan to feed into their models would require a level of transparency that may seem riskier than a possible reputation hit if it’s ever proven that a treasured first edition was destroyed in the name of advancing AI.

1.There's a huge launch crunch right now, and it will probably get worse

2.First test flight of largest all-electric aircraft used just $5 of electricity

3.Ukraine strikes major Russian rocket factory with cruise missiles

4.Meet the only known trebuchet casualty in history

5.Vulnerability giving attackers full control of Macs is under active exploitation
