Skip to content
Tech News
← Back to articles

AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale

read original more articles
Why This Matters

AI companies are increasingly purchasing and physically dissecting antique books to extract their content for training large language models, often bypassing traditional copyright restrictions through legal doctrines like fair use. This practice raises concerns about the ethics of using physical books for AI training and the potential impact on authors and publishers. As the industry shifts, the use of physical texts as 'clean' data sources highlights both opportunities and challenges in AI development and intellectual property rights.

Key Takeaways

Sign up to see the future, today Can’t-miss innovations from the bleeding edge of science and tech Email address Sign Up Thank you!

Traditionally, books have been good for two things: reading, and looking nice on a shelf.

But AI companies are interested in neither.

To those building large language models, books are nothing more than fodder to be devoured en masse before being spit out like fishbone. Often, they’re happy to use digital books — or even better, pirated digital books, as Meta has been accused of doing, and as Anthropic was forced to pay a $1.5 billion settlement to authors for also doing.

But many companies, including Anthropic, have turned to ingesting physical books instead, which they can buy countless used copies of on the cheap. According to the settled lawsuit, Anthropic used a hydraulic powered cutting machine to neatly remove the pages from the books it procured from book resellers and then scanned them using industrial-grade imaging equipment. In other words, it was literally ripping off authors’ books to train its AI.

This process took advantage of a legal concept known as first-sale doctrine, which allows a buyer to do what they want with a purchase without the original copyright holder’s say-so. And since Anthropic was turning the original physical texts into digital ones — rather than redistributing them as new copies — a judge found this to be “transformative,” and therefore protected by fair use.

Now, as 404 Media reports, this practice has become prevalent enough that even well-established book sellers are looking to cash in on the AI boom. One called ISBNdb, which boasts the “world’s largest book database,” extolls that the “world’s best AI training data is setting on a shelf,” upholding these physical texts as uncorrupted by shoddy AI writing that’s already polluted so much of the internet (and indeed, newer books).

“Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage,” it explains in an article on its website, as quoted by 404. “Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.”

Once focused on helping libraries, distributors, and book shops find and sell books, ISBNdb now helps AI companies bulk-buy anywhere between 1,000 to one million books per order, according to 404.

As an added bonus, it also promises AI companies that it’ll keep their purchases under wraps — nobody wants to end up in the spotlight like Anthropic and Meta, obviously — while clearly sounding aware about how incredibly shady the practice sounds.

... continue reading