The problem
More than seven million books downloaded from LibGen and PiLiMi entered Anthropic's central library, alongside millions of print books it lawfully bought and scanned. The court treated the tracks differently: training on lawfully acquired content was fair use, but acquiring and indefinitely retaining pirated copies was not. Final approval of the $1.5 billion settlement came on July 20, 2026, while payments and claims administration continue through September 2027.
Blast radius
~500,000 works settled; 7M+ books downloaded
Time to detect
About three years to litigation
Recovered
Partial — Settled; source files destroyed; $1.5B payable
01 — Context
Anthropic assembled books through two parallel routes. On the lawful track, it purchased millions of print books and destructively scanned them for internal use. On the other, it downloaded a bulk copy of Library Genesis in June 2021 and a copy of Pirate Library Mirror in July 2022, putting more than seven million shadow-library books into a central library. The court described that library as material retained for convenience and future optionality rather than for one justified use at acquisition.
02 — What failed
The failure sat at acquisition and retention, not in the transformative act of training. Pirated LibGen and PiLiMi copies entered a central library and were kept indefinitely for possible future uses. The court found that each use required its own justification and that convenience and optionality did not justify this separate, non-transformative use. By contrast, training on lawfully acquired content and scanning purchased print books for internal use were fair use.
03 — Symptoms
Practitioners encountered this failure through questions such as "AI company trained on pirated books," "Claude training data lawsuit," "Anthropic LibGen PiLiMi copyright," and "is training an LLM on copyrighted books fair use." The sourcing exposure did not appear as a documented model-output symptom; it surfaced in litigation filed by three authors after the shadow-library books had been retained in Anthropic's central library. The public record does not document Anthropic's internal ingestion controls or approval process.
04 — Root cause
The corpus strategy admitted and retained bulk shadow-library copies for convenience and future optionality without a lawful justification for that separate use. That standing central library was legally distinct from an active, transformative training use. The effective acquisition boundary therefore allowed pirated copies to be kept indefinitely even though the court required a separate justification for possession and retention. The internal process and controls that produced this outcome are not publicly documented.
05 — Technical explanation
The court separated what happened to book content from how copies entered and remained in the corpus. Anthropic's lawful track purchased millions of print books and scanned them for internal use. Its shadow-library track copied LibGen in June 2021 and PiLiMi in July 2022, adding more than seven million books to a central library. The court held that training on the content of lawfully acquired books was fair use, while acquiring and indefinitely retaining pirated copies in a central library was a separate, non-transformative use that was not fair use. The order therefore required a justification for acquisition and retention independent of the transformative purpose accepted for training.
06 — Contributing factors
Not publicly documented for this case.
07 — Attempted fixes
The settlement required Anthropic to destroy the original pirated files sourced from LibGen and PiLiMi. Anthropic represented that neither pirated dataset had been used in a commercially released model; the reviewed sources did not independently verify that representation. Any internal mitigation before the authors filed suit is not publicly documented.
08 — Lessons learned
Lawful use of content does not supply a lawful basis for every way copies are acquired and retained. The ruling treated transformative training, internal scanning of purchased books, and indefinite possession of pirated copies as separable uses that each required analysis. The settlement—about $3,000 per work across roughly 500,000 books—also shows how corpus-wide sourcing exposure can accumulate per work. Provenance therefore has to be resolved at acquisition, not inferred from a later training purpose.
09 — Prevention checklist
- Gate corpus admission on verified lawful-acquisition status for every external dataset.
- Record each source, acquisition date, license, and lawful basis in a provenance ledger.
- Require legal review before bulk external data is copied into a standing corpus.
- Give unverified sources a quarantine path and improperly sourced data a tested purge path.
- Set a justified retention period for source copies instead of keeping them for optional use.
References
- Order on cross-motions for summary judgment, Bartz v. Anthropic PBC, No. 3:24-cv-05417 (N.D. Cal., June 23, 2025) (opens in a new tab)
copyrightalliance.org
Court record
- District Court Issues AI Fair Use Decision (Goodwin Law client alert) (opens in a new tab)
goodwinlaw.com
Docs
-
News
- Bartz v. Anthropic Settlement: What Authors Need to Know (Authors Guild) (opens in a new tab)
authorsguild.org
Blog
- Court Grants Final Approval of $1.5 Billion Anthropic Copyright Settlement (Authors Guild) (opens in a new tab)
authorsguild.org
Blog
- Anthropic's landmark $1.5B copyright settlement is approved (TechCrunch) (opens in a new tab)
techcrunch.com
News
10 — Solutions (3)
Gate ingestion on machine-readable source provenance
Root cause addressed
This addresses bulk shadow-library material entering a retained central corpus without a lawful justification for acquisition and possession.
Make legal review a required data-pipeline stage
Root cause addressed
This addresses the need to establish a lawful basis for bulk acquisition before material is copied and retained for future use.
Limit acquisition to active needs and delete source copies promptly
Root cause addressed
This addresses indefinite retention for convenience and optionality, which the court analysed separately from transformative training.
Fixed this yourself? Add the steps that helped with Claude's training pipeline sourced millions of books from pirate sites.
Solution received
Thank you for sharing what worked. Your solution will be reviewed by a human before it appears alongside the others. It does not publish automatically.