A Court Already Said Buying Your Book and Shredding It Is Fair Use. Pirating It Cost $1.5 Billion. That Gap Is the Whole Lesson.
Anna's Archive published a call for volunteer scanners on 21 August, arguing that AI companies are buying physical books, scanning them destructively, and binning the originals, and that rare editions are disappearing into private training corpora. The post went to the top of Hacker News.
The uncomfortable part is not the practice. It is that a US federal judge already looked at exactly this and said the legal version of it is fine.
What Project Panama was
Unsealed filings in the class action against Anthropic describe an internal programme, started in early 2024, aimed at destructively scanning books at scale. Anthropic bought used copies from resellers including Better World Books, World of Books, and Zoom Books. Contractors cut the spines off, fed the pages through high-speed scanners, and disposed of the paper. Court documents describe a plan to scan somewhere between 500,000 and 2 million books under a single six-month vendor contract.
Judge William Alsup, in the Northern District of California, ruled on summary judgment that this was fair use. His reasoning on the destructive scanning was narrow and, once you read it, unsurprising: Anthropic owned the physical copies, and converting a book you own into a more convenient searchable format without making additional copies or redistributing them is a format change. He separately found that training on lawfully acquired books was "quintessentially transformative."
I am not a lawyer and this is not legal advice. But the reasoning is not exotic. It is the same logic that says you may rip a CD you bought.
The other half of the ruling
Alsup denied Anthropic summary judgment on the pirated books. The company had also acquired works through Books3, LibGen, and Pirate Library Mirror, and on that he was direct: "Anthropic had no entitlement to use pirated copies for a central library. Creating a permanent, general-purpose library was not itself a fair use."
That half ended in a settlement reported at $1.5 billion, covering roughly 500,000 works at approximately $3,000 per work, the largest copyright settlement in US history.
Put the two halves next to each other. Same company. Same models. Same training objective. The books bought and shredded produced no liability. The books downloaded produced a nine-figure cheque. Nothing in that difference is about what the model learned or what it can now generate.
The line is acquisition, not use
This is the part I think most independent writers, developers, and creators have backwards, and I had it backwards too.
The mental model in most "AI ate my work" arguments is that training on your content is the harm and copyright is the defence. The court record says otherwise. Training on a lawfully obtained copy was blessed. The liability attached to how the copy was obtained, and specifically to maintaining a permanent library of copies acquired without paying for them.
Follow that through for anything you publish. If your work is on a public web page with no access control, the acquisition is not piracy. There is no unpaid copy, no torrent, no shadow library. Under this reasoning, the strongest lever authors had was the one that only existed because Anthropic took a shortcut on procurement.
Which means the practical protection for an independent publisher is not copyright. It is whatever governs access at the moment of acquisition: a paywall, a licence agreement someone actually clicks, an API with terms attached, a contract with a distributor. Those are enforceable in ways that a "no AI training" line in your footer is not.
I find that conclusion irritating and I have not found a way around it.
What this changes about publishing independently
Three practical consequences, and none of them is "stop publishing."
Free public content is training data and always was. That is not new information, but the court record removes the ambiguity that let people hope otherwise. If you write publicly to build an audience, that trade is now explicit rather than implicit, and it is still often a good trade. I make it every week with this site.
Gated content is a different legal object. Not because gating makes it more copyrighted, but because getting past a gate requires an act, and acts leave records and create terms. The difference between "they read my public page" and "they agreed to terms and then breached them" is the entire difference between the two halves of this case.
And the moat is not the text. If your work can be reproduced by anyone who read it once, the legal system has told you fairly clearly it will not protect the reproduction. What survives is the things a corpus cannot contain: the specific data you have, your relationship with the people who read you, and being the person who publishes the update tomorrow.
Where Anna's Archive is right, and where I do not follow them
They are right that this is irreversible in a specific and unusual way. Software deprecation is annoying and recoverable. A destroyed 1930s technical manual with no other extant scan is gone. Their argument for volunteer scanning of genuinely rare material is straightforwardly good, regardless of what you think about the rest of the project.
Where I do not follow them is the framing that the destruction is the outrage. Buying a used mass-market paperback and cutting the spine off to scan it is not vandalism. Libraries and archives have done destructive scanning for decades because it produces better results than photographing bound pages. Most of what gets processed at this scale is ordinary out-of-print stock that exists in thousands of copies.
Anna's Archive also has an obvious interest in the framing. A shadow library arguing that only shadow libraries can preserve culture is making a self-serving argument, and the argument can be self-serving and correct at the same time. Both of those are true here, and I would rather say so than pick one.
What I would actually do
If you publish independently, spend twenty minutes on this and then get back to work.
Decide, explicitly, which of your work is free acquisition bait and which is gated. Most people have never made this decision and simply have everything public because that is the default. Making it deliberately changes what you write, not just where it sits.
For anything gated, check that there is an actual agreement at the boundary. A login without terms is a speed bump. A login with terms someone accepted is a contract, and this case is a very expensive demonstration that contracts and acquisition records are where the leverage lives.
And stop putting energy into "no AI training" declarations in robots.txt and footers. They are not nothing, some crawlers respect them, and respecting them is voluntary. Treat them as a preference you have registered, not a protection you have obtained.
Where this could be wrong
The largest caveat is that a district court summary judgment is not settled law. This is one judge in one district, the reasoning has not been tested on appeal in this posture, and other cases with different fact patterns are still moving. Building a publishing strategy on Alsup's reasoning as if it were a final national rule would be a mistake.
The second is that I am reading a legal outcome as a strategic signal, which is a stretch. The ruling says what was lawful for Anthropic in this case. It does not say what will be lawful for the next company, and it says nothing at all about what is decent. Plenty of things that survive summary judgment are still worth objecting to loudly, and "the court allowed it" has never been much of an argument about whether something should happen.
Author
Lukas
@lukcombinator