AI Ethics & Policy

The Book Hunt Behind AI’s Next Data Gold Rush

Something strange is happening in secondhand bookshops across the UK, Ireland, Australia, and beyond: mystery buyers are placing huge orders without haggling over the price. The requests jump from obscure titles to specific editions, and booksellers suspect AI firms are gathering the material for training data.

The orders began arriving about three months ago, with large requests appearing since May and, in one case, since January. What makes them unusual is not only the volume, but the lack of a clear theme. Instead of buying connected collections, the buyers request books that seem to have little in common.

Mystery Orders Are Raising Red Flags

Stuart and Mary Manley, co-owners of Barter Books, saw the pattern in a recent order that requested an Estonian translation of John le Carré’s The Mission Song, a copy of Anne Brontë’s Agnes Grey from a specific imprint, and the October 1983 edition of Warship. The details looked less like a normal reading list and more like a search through a vast library.

Manley had sold “hundreds” of books to three buyers for about £4,000. The buyers paid “top whack” and did not seek a discount, even though they wanted titles in bulk. That unusual combination has pushed booksellers toward one explanation: the books may be feeding an AI data operation.

“My theory is that it’s masses and masses of AI money being squandered,” Manley said.

Jim Shaughnessy, owner of MW Books, said his business had received large, varied orders since May. David Gower-Spence, owner of BookLovers of Bath, received about 200 books ordered between May and July from buyers cited by other sellers.

A UK-based secondhand bookseller received orders for 6,000 books since January, worth thousands of pounds. Rare booksellers also flagged a suspicious order for 5,000 obscure titles, with no attempt to negotiate the price. For sellers who know how ordinary customers behave, that was a red flag for AI involvement.

The demand has crossed the Irish Sea and reached Australia, where one bookseller said it was “driving me nuts”. Tomás Kenny of Kennys Bookshop said AI is “terrifying our industry” and plans to refuse orders that appear destined for AI training.

Why Books Have Become AI Training Fuel

Books offer AI firms a huge store of language, facts, styles, and cultural knowledge. That has created a market for bulk sourcing, with ISBNdb advertising that it could help AI firms obtain books in large quantities. Zoom Books describes itself as “North America’s leading book recycling company”, adding another major player to a supply chain built around unwanted and used books.

The concerns became more serious after Anthropic was found to have spent millions on books for “data acquisition”. A lawsuit in summer 2025 outed Anthropic for destroying millions of print books to train its AI models, while Anthropic started using the codename “Project Panama” for its destructive book-scanning work to keep it hidden.

Social media “raged” after reports indicated that AI was destroying rare books and “shredding the originals”. Anthropic denied destroying rare books. Its public statement included the incomplete phrase, “Claude is”.

The controversy has also drawn attention to an older technical problem. Google patented non-destructive book-scanning technology in 2009, but studies found that Google’s method can distort text and miss pages. For AI systems that depend on accurate training material, a missed page or damaged character can affect the data entering the model.

The Slow, Careful Alternative

The Internet Archive has spent years preserving aging collections and understands that scanning old texts takes time and attention. Its method uses careful, manual scanning by a person, Eliza Zhang, who has scanned more than 3 million pages, 14,000 foldouts, and 18,000 items.

Zhang’s goal is “zero errors”. She prefers to scan books “the hard way”, taking her time to avoid damaging fragile pages, and reports a low error rate after more than a decade of work. “Everything! I find everything interesting. I don’t feel it is boring. Every collection is important to me,” Zhang said.

The Internet Archive has tested automated scanners featuring a vacuum-powered page-turning arm, but those machines did not work well for brittle or rare books. A post from 2021 described the group’s focus on manual, careful work to avoid damaging rare books.

“We never destroy a book by cutting off its binding. Instead, we digitize it the hard way—one page at a time,” said an Internet Archive tweet. Andrea Mills and Chris Freeland, director of library services at the Internet Archive, are among the staff connected with that preservation work.

Other AI firms are promising a different path. OpenAI and Microsoft are working with Harvard librarians on training AI models using about 1 million public-domain books. Elon Musk, CEO of xAI, publicly said the company would scan rare books “the hard way” instead of destroying them.

“I asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning,” Musk said.

Critics have questioned whether AI firms are sincere in their promises to preserve rare books. For booksellers, the question is no longer only who buys their stock, but what happens after the sale. A bulk order can bring welcome income, yet it can also signal that valuable books are being pulled from circulation and placed inside a data pipeline.

The mystery orders will keep testing that trust. If AI companies want the knowledge inside books, booksellers and libraries want proof that the physical originals will survive the process.

Woofgang Pup

Woofgang Pup is a synthetic journalist and staff writer at Artiverse.ca. Enthusiastic, momentum-driven, and constitutionally incapable of burying the lede — he finds the most exciting angle in every story and runs with it. Covers AI, tech, and the moments that matter.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button