Loading prices...
All news
Flat vector illustration of an open book glowing bright cyan-blue with light rays and data particles bursting from its pages, symbolizing rare books being scanned to train Amazon's Nova AI models

Amazon is scanning rare books to train its Nova AI models, and destroying them

21:15 · 17.08.2026
Source: TechCrunch
0

A company that started out selling books is now buying them back only to destroy them. A 404 Media investigation placed a tracking device inside a rare book and followed it to VGT3, an Amazon facility in Las Vegas, where employees say their job consists of receiving large shipments of printed books, cutting off the spines to scan the pages faster, and feeding the scans into training data for Amazon's Nova AI models, TechCrunch reported on August 17, citing 404 Media's original investigation.

The logic behind it is straightforward even if the method is blunt: publicly available text on the open internet has largely already been scraped and used, and reusing AI-generated content to keep training larger models risks a well-documented failure mode called model collapse, where models trained on their own outputs (or each other's) gradually degrade. Rare, out-of-print books sidestep both problems — they're full of text no model has seen before, written entirely by humans, and legally acquirable simply by buying a copy rather than needing a licensing deal with a publisher. Booksellers who've dealt with these buyers suspect the operation is working systematically through ISBN numbers rather than picking titles at random, treating the used and rare book market as an untapped dataset to work through methodically.

Purchases books through commercial channels to improve the products and services customers use.

Amazon, statement to press

The practice has drawn public pushback, including from Elon Musk, who asked his own SpaceX AI team to preserve rare books and scan them without cutting the spines after the reporting circulated — a smaller, less destructive version of the same idea rather than a rejection of scanning books for training data outright. The tension isn't new so much as it's shifting targets: training-data sourcing keeps surfacing as the practical bottleneck behind flashier AI headlines, and disputes over how that data gets acquired have already produced real legal and financial consequences elsewhere in the industry, including Anthropic's roughly $1.5 billion settlement over books pirated for training rather than bought.

What makes Amazon's approach notable next to that settlement is that it's the opposite strategy on paper: buying physical copies through ordinary commercial channels sidesteps the piracy question entirely, even as it raises a different one about whether destroying the only surviving copies of out-of-print books is an acceptable cost of building a training set, especially for editions or printings that a library or archive might never get the chance to digitize once the last copy in circulation has had its spine cut off. It's a reminder that the fight over how AI models get built is playing out over more than just chatbot code and lawsuits between AI labs, the kind of dispute over who owns and controls training material that's shown up repeatedly in fights like Apple's trade secret suit against OpenAI. None of this makes the practice illegal on its own — buying a book and scanning it is a far cleaner legal position than pirating one — but it does mean the debate over training data is expanding past questions of consent and licensing into questions about preservation, since a scanned file and a physical, browsable original aren't really interchangeable once the original no longer exists anywhere to check the scan against.

None of this should be read as personalized investment advice.

Published: 21:15 · 17.08.2026
Maks

Author

Maks

Trading man

I've been interested in the cryptocurrency market for a long time, am a trader, and write articles and news about my experience and crypto in simple terms.

Comments (0)

No comments yet — be the first!