The AI Race Has Reached Tearing Apart Rare Antique Books and Not Everyone Thinks It’s Worth the Cost

The race to build bigger AI systems has expanded from scraping websites to digitizing printed books at industrial scale across the U.S. That debate has now reached rare and antique volumes, where some books are being disbound so pages can be scanned quickly, a practice that preservation specialists say can permanently alter copies that may be difficult to replace. The issue drew wider attention in June 2024 as reporting from libraries, archivists, and AI-watch groups highlighted the tension between faster data collection and long-term preservation.

Internet Archive and book digitizers are at the center of the dispute

cottonbro studio/Pexels
cottonbro studio/Pexels

The specific flashpoint involves high-speed scanning operations that remove a book’s spine so loose pages can pass through automated scanners, a method long used in commercial digitization but newly scrutinized in the AI era. Reporting published June 6, 2024, by 404 Media described concerns that old and sometimes rare books were being cut apart as companies and archives worked to expand digital text collections that can also be useful for machine learning. The Internet Archive has said in past digitization explanations that it uses both non-destructive and destructive scanning methods depending on the copy and project.

What is verified is that destructive scanning is not new, but the scale question has changed as AI developers seek larger corpora. Google Books scanned more than 40 million titles over roughly two decades, according to Google, and libraries have spent years balancing access with preservation. What remains unclear is how many antique or rare volumes have recently been disbound specifically for AI-related use, because no comprehensive national database tracks that activity and private vendors do not routinely release title-level lists.

The preservation debate is especially sensitive in research library hubs

Mikhail Nilov/Pexels
Mikhail Nilov/Pexels

The local impact is most visible in places with major rare-book collections, including New York, Boston, Washington, and the San Francisco Bay Area, where universities, archives, and specialty dealers handle older volumes every day. In those markets, librarians and conservators have said the condition, scarcity, and provenance of a specific copy matter as much as the text itself. A 19th-century book with annotations, bindings, or ownership marks can carry research value that a plain digital transcript does not preserve, according to rare-book guidance from the Library of Congress and major university libraries.

What is confirmed is that many institutions already have preservation rules that limit destructive scanning for unique or fragile items. What is not yet known is whether AI demand will push more scanning work toward copies sourced from private resellers rather than public libraries, where oversight can be stricter. Dealers and archivists interviewed in 2024 coverage said inexpensive older books can be attractive for bulk digitization because they are easier to buy outright, even when surviving copies are not abundant.

The pressure comes from AI training demand and the economics of speed

Google DeepMind/Pexels
Google DeepMind/Pexels

The reason this is happening is straightforward: text remains one of the core raw materials for large language models, and scanning bound books by hand is slower and more expensive than cutting them for feed scanners. Preservation experts have said non-destructive capture can take significantly longer per volume, while bulk operators are judged on throughput. Court fights involving OpenAI, Anthropic, Meta, and other AI companies during 2023 and 2024 also put new attention on where training text comes from and whether licensed or scanned book corpora may become more valuable.

For readers and researchers, the practical takeaway is that digitized access may expand while some original physical copies disappear from normal handling or lose part of their artifact value. Libraries have not announced any nationwide policy shift tied specifically to AI as of June 2024, and no federal rule bars owners from disbinding books they legally possess. The broader policy question, according to archivists and digital preservation specialists quoted in 2024 reporting, is whether faster access should come at the cost of irreplaceable physical history.

Similar Posts