If you advertised here, this is how many people would have seen your ad in just one day. Click to learn more.
Potential Views
500,001
LEARN MORE
Why is AI destroying millions of books: Inside the destructive scanning industry.

Why is AI destroying millions of books? The legal loophole behind modern-day book burning

AI is destroying millions of books because the technology industry has discovered that legally purchasing physical copies, digitising them and retaining private digital versions can provide enormous quantities of high-quality training data while exploiting unresolved boundaries in copyright law.

The practice, commonly called destructive scanning, involves removing a book’s binding, scanning its pages at industrial speed and disposing of the physical volume. The controversy became particularly acute after court documents revealed Anthropic’s Project Panama, an effort described internally as seeking to “destructively scan all the books in the world”.

The central issue is not simply whether a company may own a book. Under US law, ownership of a physical copy and ownership of the copyright are separate legal interests. The deeper question is whether purchasing a lawful physical copy, reproducing its entire contents digitally, destroying the purchased object and retaining the resulting digital archive for commercial AI development falls within fair use. A 2025 ruling in Bartz v Anthropic found that Anthropic’s use of lawfully acquired books for AI training was fair use, while finding that its use of pirated books to build a central library was not.

Key Takeaways

  • AI companies need large quantities of high-quality human-authored material for model training.
  • Destructive scanning converts physical books into private machine-readable digital archives.
  • Fair use does not automatically authorise copying every book for every purpose.
  • The first sale doctrine governs ownership of particular copies, not ownership of copyright.
  • Current law was developed before private AI libraries could contain millions of digitised books.

The new meaning of “book burning

The phrase “modern-day book burning” is deliberately provocative, but it describes a real cultural concern arising from a technologically new process. Traditional book burning was an attempt to destroy access to knowledge. Destructive scanning has a different immediate objective: extract the information from the physical object before disposing of the object itself.

The distinction matters. The books are not necessarily being destroyed because their ideas are unwanted. They are being purchased because their contents are valuable. Once the text has been converted into machine-readable data, the physical artefact can become economically redundant to the company performing the scan.

That distinction makes the practice simultaneously less dramatic and more consequential than traditional book destruction. The objective is not necessarily to prevent humanity from reading a particular book. It is to transform a physical cultural object into proprietary computational infrastructure.

Anthropic provides the clearest publicly documented example. Court disclosures revealed that the company spent millions of dollars acquiring millions of printed books, often used copies, and employed contractors to remove bindings, prepare the pages and scan them into digital files. The physical originals were then discarded.

The result is a striking inversion of the traditional economics of publishing. A book once had value because somebody wanted to read it. In the AI economy, the same book can have value because a machine wants to ingest its language patterns, information, structure and relationships.

Why AI wants books published before 2022

There is a technical reason older books have become particularly attractive.

The public internet is increasingly saturated with machine-generated material. Search engines, social platforms, websites and publishing systems now contain enormous quantities of text produced or modified by generative AI. Training future models indiscriminately on material generated by previous models presents a potential quality problem.

Research into model collapse has identified circumstances in which recursive training on synthetic data can reduce diversity and distort the distribution of information represented by a model. The problem is not that every AI-generated sentence is inherently unusable. Rather, repeated generations of synthetic material can progressively amplify errors, biases and statistical artefacts.

Books published before the mass adoption of generative AI therefore possess an important property: they constitute a large reservoir of predominantly human-produced language created before today’s AI-content explosion.

This does not mean every book published before 2022 is necessarily human-written, nor does it mean every post-2022 publication is AI-generated. It means that older books provide a comparatively clean historical corpus whose provenance is easier to establish.

For an AI company, millions of professionally edited books represent something considerably more valuable than random internet text. They contain long-form reasoning, narrative structure, specialist terminology, historical knowledge, dialogue, argumentation and carefully edited language.

The commercial incentive is therefore enormous.

App Development with Flutter and Dart
Be able to build any Android or iOS app you want based on Dart and Flutter. Save time, money and your sanity when you learn Dart and Flutter in our budget-friendly, jargon-free video course now. Don’t let a lack of technical knowledge and coding stand in your way. In this video course you’ll learn what building an app entails and how to make a user-friendly app without it being frustrating, time-consuming or costly. The exciting moment when you start running your first Flutter app starts now. Recommended Study Time: 30 hours.

The technology makes destruction economically rational

Destructive scanning is not a technological necessity. It is an efficiency decision.

A bound book can be digitised without damaging it using specialised overhead scanners, robotic page-turning systems or other non-destructive techniques. Such systems are particularly important for libraries, archives and institutions preserving rare or fragile material.

Destructive scanning takes a different approach. The binding is removed, pages are separated and the resulting loose sheets can be processed through high-speed document scanners. At industrial scale, this can dramatically increase throughput while reducing the labour associated with carefully preserving each volume.

The physical book consequently becomes an input rather than an object of preservation.

That distinction is critical. A library normally digitises a book because it wants to preserve and provide access to the work. An AI company can digitise a book because it wants the information contained within it as a computational resource.

The economic objective is therefore radically different.

The legal problem begins with copyright

US copyright law gives copyright owners exclusive rights including reproduction, distribution and the preparation of derivative works, subject to statutory limitations and exceptions.

A company buying a physical book does not automatically purchase those copyright rights.

That principle is explicitly recognised by 17 U.S.C. §202, which distinguishes ownership of copyright from ownership of the material object in which a copyrighted work is embodied. Buying the physical object does not transfer copyright in the work.

This is where fair use becomes central.

Section 107 of the Copyright Act establishes fair use as a limitation on copyright infringement. Courts consider the purpose and character of the use, the nature of the copyrighted work, the amount used and the effect on the potential market. Commercial use can be relevant, although it is not automatically disqualifying.

The law was designed to accommodate socially valuable uses of copyrighted material without requiring permission for every quotation, criticism, scholarly use or research activity.

The problem is that artificial intelligence can operate at a scale that previous copyright cases never contemplated.

A researcher photocopying several pages and a corporation creating a private digital corpus containing millions of complete books are technically performing the same fundamental act, reproduction, but at dramatically different scales and for dramatically different economic purposes.

Data Engineering Courses
Data Engineering Courses
Data engineering courses can help you learn data modeling, ETL (extract, transform, load) processes, and data warehousing techniques. You can build skills in data pipeline construction, database management, and ensuring data quality and integrity. Many courses introduce tools like Apache Spark, Hadoop, and SQL, that support processing large datasets and optimizing data workflows. You’ll also explore cloud platforms such as AWS and Azure, which facilitate scalable data solutions and enhance your ability to manage data in various environments.

Google Books helped create the legal foundation

The AI industry’s current legal position did not emerge from nowhere.

The landmark Authors Guild v. Google litigation concerned Google’s scanning of millions of copyrighted books. The Second Circuit concluded that Google’s digitisation and search functions constituted fair use, and the US Supreme Court declined to review the decision in 2016.

Google’s project was important because the company was not simply reproducing books for people to read. Its system enabled users to search the corpus and obtain limited information about books, including snippets, while creating a searchable index.

That distinction helped establish a powerful concept in copyright jurisprudence: copying an entire work can, in particular circumstances, be legally permissible when the purpose of the copying is sufficiently transformative and the resulting use does not function as a conventional substitute for the original.

AI companies now operate in that legal territory.

They argue that training a model is not equivalent to publishing millions of books. The model does not normally provide users with a conventional digital library containing every source text. Instead, the company argues, the books are transformed into statistical information used to construct a new computational system.

Judge William Alsup’s 2025 ruling in Bartz v. Anthropic accepted a version of that reasoning. The judge concluded that Anthropic’s use of lawfully acquired books to train its models was fair use and described the transformation as exceptionally substantial.

That ruling is enormously important, but it does not mean “AI can legally copy books”.

It means a particular court found a particular use, under particular facts, to be fair use.

The first sale doctrine is not a copyright licence

The first sale doctrine is frequently misunderstood in this debate.

Section 109 of the Copyright Act generally allows the lawful owner of a particular copy of a copyrighted work to sell or otherwise dispose of possession of that particular copy without obtaining permission from the copyright owner.

Its historical purpose is straightforward. Copyright law grants copyright owners control over distribution, but that control cannot continue indefinitely over every physical copy once the copyright owner has lawfully transferred ownership.

Without the doctrine, someone buying a second-hand book could theoretically face copyright restrictions every time they attempted to resell or dispose of it.

The doctrine therefore protects the ordinary circulation of physical goods.

What it does not do is transfer copyright in the underlying work.

This distinction is essential to understanding the AI book controversy.

Buying a book legitimately means the buyer owns that copy. It does not ordinarily mean the buyer can reproduce the entire book, upload the reproduction, distribute millions of digital copies or establish a commercial database containing the work.

The controversial legal argument arises because AI companies can combine lawful ownership of a physical copy with a separate fair-use argument for reproducing its contents.

The first sale doctrine supplies the lawful acquisition. Fair use supplies the argument for digitisation.

That combination is far more consequential than either doctrine considered independently.

AWS Certified Solutions Architect Associate SAA C03 Specialization
AWS Certified Solutions Architect Associate (SAA-C03) Specialization
Master AWS Cloud Architecture and Get Certified. Gain hands-on experience with AWS, learning cloud architecture, security, and cost optimization.

Why destroying the physical copy is an issue

This is where the modern controversy becomes particularly unusual.

In Bartz, the court accepted that Anthropic’s digitisation of lawfully acquired books and subsequent destruction of the physical copies could constitute a lawful one-to-one format conversion for the purpose at issue. The court separately rejected Anthropic’s acquisition of pirated books for its central library as fair use.

The legal significance is therefore not that destroying a book magically makes copying lawful.

It is that destruction can support an argument that the company has converted one lawfully acquired copy into one digital copy rather than creating an additional conventional copy for distribution.

That is a critical distinction.

The physical book disappears. The digital representation remains.

From the company’s perspective, the information survives in a form vastly more useful to an AI system. From a cultural perspective, however, one physical manifestation of the work has disappeared from circulation.

That becomes particularly troubling when the book is out of print, obscure or difficult to replace.

When legal compliance can still create a cultural problem

This is the point at which the phrase “book burning” becomes analytically useful.

Something can be legally permissible and culturally destructive at the same time.

A corporation may lawfully purchase a used book from a bookseller. The bookseller may lawfully sell it. The corporation may lawfully own the physical copy. A court may ultimately determine that the digitisation was fair use.

Yet the cumulative result can still be a reduction in the number of physical copies circulating through society.

For common mass-market books, this may have little practical significance. Millions of copies may remain available.

For obscure academic works, specialist technical books, regional publications, small-print literary works or out-of-print titles, the calculation is different.

The disappearance of one copy may not matter.

The disappearance of thousands of copies across multiple markets can matter enormously.

The most important issue is therefore not whether AI companies are literally erasing every book they acquire. The evidence does not support that sweeping claim. Nor does the available evidence establish that millions of unique, irreplaceable rare books have already been removed from existence.

The documented concern is more precise: industrial-scale digitisation can remove physical copies from circulation while concentrating the resulting digital information inside private corporate systems.

Google IT Support Professional Certificate
Google IT Support Professional Certificate
The launchpad to a career in IT. This program is designed to take beginner learners to job readiness in about three-to-six months.

The private library is the bigger issue

A public digital archive could produce an extraordinary social benefit.

Imagine a searchable, preservation-grade database containing millions of books, made available to libraries, researchers, universities and the public under appropriate copyright controls.

That would represent a historic expansion of access to human knowledge.

The AI industry’s private databases are different.

The digitised material can become part of proprietary infrastructure used to develop commercially valuable models. The resulting corpus may be inaccessible to researchers, libraries, authors and the public.

The cultural value of digitisation is therefore being separated from the economic value of digitisation.

The books become inputs into private computational systems rather than components of a shared cultural archive.

That distinction deserves substantially more attention from legislators.

The law is struggling to catch up with the technology

This is fundamentally a problem of technological acceleration.

Copyright law developed around identifiable acts of publication, copying, distribution and commercial exploitation. Digital technology subsequently made copying dramatically cheaper. Artificial intelligence has now made the economic value of copying potentially enormous.

The legal system is being asked to apply doctrines developed for earlier technological environments to corporations capable of processing billions of pages.

The scale changes the consequences.

A legal doctrine that produces an acceptable result when applied to thousands of works can produce a profoundly different result when applied to millions.

The commercial incentive is also unusual. AI companies compete intensely for better training data. Better data can produce better models. Better models can generate enormous commercial valuations.

That creates a powerful incentive to acquire data quickly, cheaply and comprehensively.

The litigation surrounding AI demonstrates that this is not merely theoretical. In July 2026, Hachette Book Group, Cengage Learning, Elsevier and author Scott Turow filed a new federal lawsuit accusing Google of copying millions of textual works to develop Gemini. Those allegations remain allegations and have not been adjudicated.

Similar litigation has been brought against other major AI companies, demonstrating that the boundaries of training-data law remain unsettled.

IBM Data Engineering Professional Certificate
IBM Data Engineering Professional Certificate
Prepare for a career as a Data Engineer. Build job-ready skills – and must-have AI skills – for an in-demand career. Earn a credential from IBM. No prior experience required.

This is why the first sale doctrine needs modern scrutiny

The first sale doctrine itself is not the villain.

Its original function remains important. People should be able to buy, sell, lend, donate and dispose of physical books without seeking permission from copyright holders every time ownership changes.

The difficulty arises when industrial technology allows a corporation to use lawful ownership of physical copies as the starting point for creating a massive private digital corpus.

The doctrine was never designed around that economic model.

Neither was copyright law designed around a world in which a company could purchase millions of books, convert them into machine-readable data, discard the physical copies and use the resulting information to develop an artificial intelligence system potentially worth hundreds of billions of dollars.

Calling this a malicious “use” of first sale doctrine is therefore better understood as a criticism of the interaction between legal doctrines rather than a claim that Section 109 itself expressly authorises AI training.

It does not.

The unresolved issue is how first sale, fair use, reproduction rights and emerging AI technologies should interact when applied at unprecedented scale.

The real danger is not that every book will disappear

The strongest argument against AI book destruction is not that civilisation’s entire literary heritage is about to vanish.

It is that a legal system designed to balance competing interests can produce unintended incentives when technology changes the economics of those interests.

If buying a physical book and destroying it after digitisation is cheaper and legally safer than negotiating a licence, companies have an incentive to buy books rather than licence them.

If private digital archives provide competitive advantages, companies have an incentive to keep them secret.

If the law treats training as sufficiently transformative, companies have an incentive to maximise the volume of material they can acquire.

And if competitors are doing the same thing, every company has an incentive to move faster.

That is how a system can incentivise behaviour that is individually rational while producing a collectively troubling result.

The central question is therefore no longer merely whether AI training is “fair use”.

The more important question is whether a fair-use framework designed for a different technological age remains capable of protecting authors, publishers, libraries and cultural heritage when artificial intelligence can transform millions of copyrighted works into private industrial infrastructure.

Samsung Galaxy Book6 Ultra
  • BIG POWER: Create without compromise on a powerhouse laptop for AI and productivity with the Galaxy Book6 Ultra
  • INTEL ARC GRAPHICS TO TAKE YOUR IDEAS FURTHER: This powerful PC elevates every frame, shot and render of graphics-intensive projects like 3D animations and AI generated videos
  • DYNAMIC AMOLED 2X DISPLAY: With the power to show rich and vibrant colours and a refresh rate of up to 120Hz, all your creative projects come alive in vivid clarity
  • SIX SPEAKERS: Tuned with Dolby ATMOS, the six-speaker system is a first for Galaxy PC computers
  • SLIM & LIGHT LAPTOP: Experience a premium two-tone keyboard, a slimmer hinge and a lightweight, symmetrical frame that’s easy to carry
  • KEEPS COOL: New and improved vapour chamber and redesigned fans and vents that work together to keep your laptop cooler

The future of copyright may depend on the answer

The United States has already demonstrated that courts can recognise transformative technological uses of copyrighted material. Authors Guild v. Google established one important precedent, while Bartz v. Anthropic has extended the debate into generative AI.

But neither case settles the entire AI copyright question.

The distinction between lawful acquisition and unlawful acquisition remains crucial. The distinction between training and maintaining a permanent private library remains crucial. The distinction between creating a new technological tool and commercially substituting for copyrighted works remains crucial.

And the distinction between legal ownership and cultural preservation may become increasingly important.

AI development does not inherently require the destruction of books. Digitisation can preserve books. Libraries have demonstrated this for decades. Modern scanning technology can capture enormous quantities of information without necessarily destroying the source material.

The fact that destructive scanning can nevertheless be economically attractive demonstrates precisely what is at stake.

The technology is not forcing society to destroy books.

The economics are rewarding companies for extracting their information as efficiently as possible.

That is why why is AI destroying millions of books is ultimately a copyright question, a technology question and a cultural preservation question at the same time.

The most important lesson from the current controversy is not that artificial intelligence is literally conducting a campaign to erase human knowledge. The evidence does not support that conclusion. The more serious and defensible conclusion is that AI companies have discovered how existing copyright doctrines can, under certain circumstances, permit them to convert enormous quantities of privately owned physical literature into proprietary digital resources.

That is a profound development.

Copyright law has always attempted to balance technological progress against the interests of creators and the public. The printing press, photocopier, recording industry, software, the internet and digital publishing each forced the law to reconsider where that balance should lie.

Artificial intelligence is creating the next confrontation.

The question facing legislators and courts is whether a doctrine intended to protect legitimate research and technological innovation can continue to perform that function when the world’s most powerful technology companies can use it at industrial scale.

If the answer is no, the solution is not to abolish fair use or the first sale doctrine. It is to modernise the rules around them so that a legal mechanism intended to promote access to knowledge does not inadvertently encourage the permanent removal of physical cultural objects and the concentration of their digitised contents inside private corporate archives.

That is the real modern-day book-burning problem.

Not that the words disappear.

It is that the physical books can disappear while the knowledge becomes private property in everything but name.

Recent Articles

When you buy something through our retail links, we may earn commission and the retailer may receive certain auditable data for accounting purposes.

WhatsApp Channel Follow Sweet TnT Magazine on WhatsApp

Amazon eGift card

Every month in 2026 we will be giving away one Amazon eGift Card. To qualify subscribe to our newsletter.

You may also like:

How to self-publish a book: A complete beginner’s guide

Manuscript editing: The foundation of a flawless self-published book

The art of storytelling: Why personal narratives still matter

The daily drip: What happens when you write a little every day?

The power of words: Exploring poetry in the modern world

The art of letter writing: Connecting in a digital world

The key to speed writing and how to write any assignment fast

Student Writing Tips: Make best assignments fast

7 Effective ways to improve your academic writing skills

5 Effective ways rhyming books help to improve reading skills

No more confusion! A guide to popular essay formats

6 Best resources for parents to help improve child’s reading skills

@sweettntmagazine

Discover more from Sweet TnT Magazine

Subscribe to get the latest posts sent to your email.

About Jevan Soyer

Jevan Soyer draws from a multifaceted career spanning the hospitality, tourism, education, sales, marketing and construction industries, he brings a methodical and disciplined approach to digital media. A father of two sons, marketing manager and content creator for Sweet TnT Magazine, Study Zone Institute, co-author and editor of Sweet TnT Short Stories and Sweet TnT 100 West Indian Recipes,Soyer specialises in documenting the biodiversity and cultural heritage of Trinidad and Tobago for a global audience. For editorial submissions, advertising opportunities, or to request a media kit, please contact the team directly at contact@sweettntmagazine.com.

Check Also

The history and cultural power of tabanca in the Caribbean.

Tabanca: The heartache that defines Trinidad and Tobago

Tabanca is Trinidad and Tobago’s defining expression for emotional heartbreak, longing, and psychological yearning. More …

The literary works of Eric Williams: History, politics and Caribbean identity.

Eric Williams: The complete literary works of Trinidad and Tobago’s scholar-statesman

Eric Williams remains the most influential Caribbean historian and political intellectual of the twentieth century, …

Leave a Reply

Discover more from Sweet TnT Magazine

Subscribe now to keep reading and get access to the full archive.

Continue reading