2026-09-21 16:09:12

One of the largest digital archives spoke out against scanning books for AI

One of the largest digital archives spoke out against scanning books for AI

The open online archive Anna’s Archive is calling on volunteers around the world to help digitize and preserve rare and limited-edition books. The archive's specialists are alarmed: too often, printed publications are simply "fed" to AI and then destroyed.

A striking example is the recent courtroom defeat of Anthropic, which was found guilty of copyright infringement in 2024 (the company had used "pirated" books to train its AI models) and ordered to pay nearly one and a half billion dollars. The court did clarify, however, that using existing works to train artificial intelligence models may be considered fair use and not a violation of copyright. This ruling has given a green light to such "training" practices: the market is showing active buying up of rare, limited-print, and otherwise "illiquid" books.

Books are becoming raw material for AI

Increasingly, printed books are becoming not a direct source of knowledge, but «fodder» for AI and other digital content generators. Having processed books and trained on them, neural networks produce new texts, without always being accountable for data quality. This is precisely what makes rare books so valuable: they are full of examples of unusual structure, terminology, and specialized vocabulary — all of it carefully edited, which is something one encounters ever less frequently even on the modern internet. In short, as a «teacher» and a «source,» such publications are a genuine treasure.

Yet their subsequent fate looks truly barbaric. Scanning a book takes minutes, but for the convenience of the machine, the spine is often cut off and the binding destroyed — and in this «disassembled» state the book is no longer of use to anyone and is sent to the scrap heap. In the case of truly rare editions, companies even see this as an advantage: supposedly, other corporations will not be able to train their AI models on that same book!

In the long run, however, this may cause a significant impoverishment of libraries, private collections, and secondhand bookshops.

Anna’s Archive is looking for volunteers for digitization

And from all over the world! If 10 million people scan at least one book each, it will be possible to preserve 10 million unique copies. If someone is willing to undertake mass digitization, the archive is even prepared to discuss reimbursing scanning costs.

Anna’s Archive currently has more than 61.6 million books and around 95 600 000 articles. At the same time, there are certain objections to such digital archives: rights holders do not always agree to distributing digital copies of their works so widely, viewing it as piracy.

Your comment / review / question
There are no comments here yet
Your comment / review
If you have a question, write it, we will try to answer
* - Field is mandatory
Chat with us, we are online!

Request a call

By submitting a request, you accept the conditions Privacy Policy