Anthropic and the Stolen Library
To build Claude, Anthropic needed books. Not a few thousand, but millions.
The simplest way to obtain them was to buy them. The fastest was to download them. Anthropic did both, and that apparently mundane distinction produced one of the most interesting copyright decisions AI has yet brought us.
Between 2021 and 2022, Anthropic assembled a digital library of almost unimaginable size. Federal Judge William Alsup found that it downloaded 196,640 books from Books3, at least five million from Library Genesis, and at least another two million from Pirate Library Mirror. More than seven million copies from pirate libraries: not links, extracts or text fragments incidentally collected online, but complete PDF, EPUB or TXT books, often accompanied by bibliographic details.
Anthropic later changed methods. It began buying enormous quantities of printed books, dismantling them, cutting off their bindings, scanning each page and destroying the originals after making digital copies.
At the end of either process, its servers contained files holding books. Technically, they might seem the same. Legally, they were not.
When authors sued, something less intuitive than the predictable clash between writers and AI occurred. In June 2025, Judge Alsup held that using the books to train Claude could constitute fair use. He considered training highly transformative: the model was not using works to republish them as substitutes for the originals, but to learn linguistic relationships and generate something new.
Alsup compared the process to a reader aspiring to become a writer. We read authors, absorb structures, words, ideas and arguments, then produce something of our own. Copyright protects the work; it does not give its author a monopoly over everything others can learn by reading it.
So far, an enormous victory for Anthropic and the AI industry more broadly.
But in the same decision, the judge asked a much simpler question: where did the books Claude could read come from?
Here the story changes completely.
A lawful subsequent use does not necessarily make the way the work was acquired lawful.
An analogy helps.
I can buy a novel, read it, study its style and use what I learn to write my own. Whether my new novel infringes the earlier author’s copyright is one question.
But if, instead of buying the book, I steal it from a bookshop, the lawfulness of what I later write does not erase the earlier problem.
Stealing a book does not become lawful merely because reading it is lawful.
That is precisely the distinction applied to Anthropic. It could argue that training Claude on books was fair use; it could not use that conclusion to turn the earlier download of millions of pirated copies retroactively into fair use.
Alsup was particularly clear: US copyright law contains no special exception for AI companies. Significantly, he treated differently even two operations producing the same material result. Digitising a lawfully purchased printed book to replace it with a digital copy in an internal library could be fair use; downloading that same book from LibGen to build the same library was not.
The distinction is only superficially subtle.
The final file is identical. Its words are identical. The algorithm that later reads them is identical. What changes is the copy’s legal history.
And that history cost Anthropic dearly.
On 20 July 2026, the federal court gave final approval to a settlement under which Anthropic will pay $1.5 billion to resolve the class action concerning its pirate library. It is the largest known US copyright class-action settlement.
The action covered 482,460 works, giving a theoretical value of roughly $3,100 per work before costs and fees. The court noted that this exceeded four times the Copyright Act’s minimum for ordinary infringement and fifteen times its minimum for innocent infringement. Anthropic must also destroy the relevant LibGen and Pirate Library Mirror copies, except those retained to comply with litigation obligations.
One misunderstanding must be avoided, because it is what makes this case interesting: Anthropic is not paying $1.5 billion because a judge ruled that training AI on books is unlawful.
Almost the opposite.
On training, Anthropic had secured a hugely important victory. The unresolved problem concerned how it built its library.
One of the world’s largest AI disputes was therefore resolved by separating two verbs often conflated in public debate: using and obtaining.
From books to songs
The distinction becomes still more interesting if we move from books to music.
Broadly speaking, Suno and Udio operate on similar logic: a model encounters vast amounts of music and learns harmonic, rhythmic, timbral and stylistic structures from which it can later generate new tracks.
Again, debate quickly focuses on the result: if I ask AI for a new song, how similar must it be to an existing song to infringe copyright? May it imitate a genre? A singing style? A particular sound?
These are important questions. But Anthropic suggests another that logically comes first: which recordings were copied to let the machine learn all this?
This is not theoretical.
In 2024, Universal, Sony and Warner sued Suno and Udio, alleging unauthorised copying of protected recordings to train their models. Since then, the picture has become more interesting: Warner and Universal reached agreements with Udio, and Warner with Suno, turning some litigation into licensing relationships designed to build authorised music models.
The story is far from over.
In recent days, Sony has opened a new front against Udio, alleging that audio fingerprinting identified over 30,000 of its recordings used in the system. Naturally, that allegation must be proved in court, and the legal framework differs from Anthropic’s. Yet the underlying question is strikingly similar.
Suppose a court eventually held that letting AI “listen” to millions of songs to learn to compose new music was, under certain conditions, transformative use.
The matter would not necessarily end there.
We would still have to ask how those millions of songs entered the machine’s record collection.
That may be the most interesting lesson of Bartz v. Anthropic. AI forces us to discuss enormously sophisticated technological problems, but does not erase the elementary legal questions that precede them.
We can debate at length whether a machine truly “reads” a book or “listens” to a song, whether learning differs from copying, or whether artificial learning should be treated differently from human learning. Law will grapple with these questions for years, and the United States and Europe may reach very different answers.
Anthropic shows that, in AI law, the decisive question is not just what a model did with a protected work, but how that work entered the system.
Perhaps that is the most interesting part of the story. After more than seven million books, infrastructure capable of turning them into training material and a $1.5 billion settlement, the problem reaches a deeper question: who owns the knowledge contained in a work?
Copyright protects the work, not everything a reader can learn from it. Authors can claim their text, music and creative expression; claiming ownership of ideas, structures, information or abilities that others develop after reading or listening is much harder.
This is where AI puts pressure on categories designed with humans in mind. If learning from a work does not necessarily mean copying it, we must still determine the conditions under which that knowledge may be acquired.
Perhaps this is the real boundary Anthropic reveals: knowledge may belong to nobody, but the route by which it is acquired is not therefore legally irrelevant.
Paolo Fortina · Originally published on LinkedIn on 24 July 2026. Read the original


Comments