GKRootWire
AI Google Adds 'Preferred Source' Button to Help Publishers Fight AI Traffic LossesGadgets Linkdaze Launches a Smart Calendar Aimed at Running Your Whole HouseholdSecurity Popular Rust Crate arrayref Hijacked to Spread Infostealer MalwareCloud & Sysadmin GitHub Details Cause of August 17 Outage, Outlines Reliability FixesDev Tools Show HN: 'Huzzah' Proposes a Fresh Take on AI-Assisted CodingCloud & Sysadmin The Weird Science of Cooling Data Centers With UrineAI Google Adds 'Preferred Source' Button to Help Publishers Fight AI Traffic LossesGadgets Linkdaze Launches a Smart Calendar Aimed at Running Your Whole HouseholdSecurity Popular Rust Crate arrayref Hijacked to Spread Infostealer MalwareCloud & Sysadmin GitHub Details Cause of August 17 Outage, Outlines Reliability FixesDev Tools Show HN: 'Huzzah' Proposes a Fresh Take on AI-Assisted CodingCloud & Sysadmin The Weird Science of Cooling Data Centers With Urine
AI

AI Training Push Is Reportedly Destroying Rare Physical Books

A shadow library project is calling for mass digitization efforts before more original volumes are lost to bulk scanning processes.

A blog post from the operators of a shadow library argues that some AI companies acquiring large book collections for training data are using destructive scanning methods - slicing bindings apart to feed pages through automated scanners - and then discarding the physical originals afterward.

The post frames this as an urgent preservation issue, particularly for rare or out-of-print books where the scanned copy may become the only surviving version. It calls for coordinated efforts to digitize vulnerable texts more carefully, or at least in parallel with any destructive processes, before irreplaceable originals disappear.

The discussion on Hacker News centered on how AI training data pipelines treat physical media as disposable input rather than artifacts worth preserving, and what obligations (if any) buyers of bulk book lots have to libraries or archives.

Why it matters: AI labs are quietly becoming major consumers of physical books as training data, but the digitization methods driving that appetite operate with none of the preservation standards libraries use. If scanning-and-shredding becomes standard practice, some texts could lose their only physical record - a real cultural cost baked into the AI supply chain that gets little scrutiny compared to copyright debates.

Sources: Hacker News