GKRootWire
Cloud & Sysadmin Microsoft Confirms Preview Update Wipes Out Desktop SettingsAI Nvidia to Acquire Hugging Face for $12.9 BillionDev Tools A Deep Dive Into Intrusive Linked ListsGadgets DJI's Romo 2 Robovac Adds Local-Only Mode After Privacy ScareAI Nvidia Reportedly Moves to Acquire Hugging FaceAI Anthropic Launches Claude Tools for AI Shopping AgentsCloud & Sysadmin Microsoft Confirms Preview Update Wipes Out Desktop SettingsAI Nvidia to Acquire Hugging Face for $12.9 BillionDev Tools A Deep Dive Into Intrusive Linked ListsGadgets DJI's Romo 2 Robovac Adds Local-Only Mode After Privacy ScareAI Nvidia Reportedly Moves to Acquire Hugging FaceAI Anthropic Launches Claude Tools for AI Shopping Agents
AI

Amazon Is Reportedly Scanning and Destroying Rare Books to Feed AI Models

With the open internet largely tapped out, Amazon is turning to physical rare books as fresh training data for its large language models.

Amazon has reportedly been acquiring rare and out-of-print books, scanning their contents, and destroying the physical copies afterward, all to gather fresh text for training its AI models. The logic is straightforward: most publicly available online text has already been scraped and used, so unique, never-digitized books offer a rare source of novel language data.

Critics argue this process is troubling because it permanently destroys physical artifacts, some of which may be irreplaceable, in service of feeding a data-hungry AI pipeline. It also raises questions about transparency, since there's no public catalog of what's being scanned or destroyed, and no guarantee the content will ever benefit anyone beyond Amazon's own models.

The practice highlights just how strained the supply of fresh training text has become for large AI labs.

Why it matters: As internet-scraped data dries up, AI companies are increasingly hunting for untapped text sources, and physical archives are apparently fair game even at the cost of destroying originals. This raises real preservation and accountability concerns, especially if there's no public record of what's lost in the process.

Sources: TechCrunch