Skip to content
Submit a case
AFL-0004 Incident Fine-tuning & Training

Claude's training pipeline sourced millions of books from pirate sites

Claude Legal Training data provenance

Incident

June 2021 – July 2026 (acquisition through final judgment)

Published

11 Aug 2026

Sources

6

Verification

Court-verified

The problem

More than seven million books downloaded from LibGen and PiLiMi entered Anthropic's central library, alongside millions of print books it lawfully bought and scanned. The court treated the tracks differently: training on lawfully acquired content was fair use, but acquiring and indefinitely retaining pirated copies was not. Final approval of the $1.5 billion settlement came on July 20, 2026, while payments and claims administration continue through September 2027.

Blast radius

~500,000 works settled; 7M+ books downloaded

Time to detect

About three years to litigation

Recovered

Partial — Settled; source files destroyed; $1.5B payable

01 — Context

Anthropic assembled books through two parallel routes. On the lawful track, it purchased millions of print books and destructively scanned them for internal use. On the other, it downloaded a bulk copy of Library Genesis in June 2021 and a copy of Pirate Library Mirror in July 2022, putting more than seven million shadow-library books into a central library. The court described that library as material retained for convenience and future optionality rather than for one justified use at acquisition.

02 — What failed

The failure sat at acquisition and retention, not in the transformative act of training. Pirated LibGen and PiLiMi copies entered a central library and were kept indefinitely for possible future uses. The court found that each use required its own justification and that convenience and optionality did not justify this separate, non-transformative use. By contrast, training on lawfully acquired content and scanning purchased print books for internal use were fair use.

03 — Symptoms

Practitioners encountered this failure through questions such as "AI company trained on pirated books," "Claude training data lawsuit," "Anthropic LibGen PiLiMi copyright," and "is training an LLM on copyrighted books fair use." The sourcing exposure did not appear as a documented model-output symptom; it surfaced in litigation filed by three authors after the shadow-library books had been retained in Anthropic's central library. The public record does not document Anthropic's internal ingestion controls or approval process.

04 — Root cause

The corpus strategy admitted and retained bulk shadow-library copies for convenience and future optionality without a lawful justification for that separate use. That standing central library was legally distinct from an active, transformative training use. The effective acquisition boundary therefore allowed pirated copies to be kept indefinitely even though the court required a separate justification for possession and retention. The internal process and controls that produced this outcome are not publicly documented.

05 — Technical explanation

The court separated what happened to book content from how copies entered and remained in the corpus. Anthropic's lawful track purchased millions of print books and scanned them for internal use. Its shadow-library track copied LibGen in June 2021 and PiLiMi in July 2022, adding more than seven million books to a central library. The court held that training on the content of lawfully acquired books was fair use, while acquiring and indefinitely retaining pirated copies in a central library was a separate, non-transformative use that was not fair use. The order therefore required a justification for acquisition and retention independent of the transformative purpose accepted for training.

06 — Contributing factors

Not publicly documented for this case.

07 — Attempted fixes

The settlement required Anthropic to destroy the original pirated files sourced from LibGen and PiLiMi. Anthropic represented that neither pirated dataset had been used in a commercially released model; the reviewed sources did not independently verify that representation. Any internal mitigation before the authors filed suit is not publicly documented.

08 — Lessons learned

Lawful use of content does not supply a lawful basis for every way copies are acquired and retained. The ruling treated transformative training, internal scanning of purchased books, and indefinite possession of pirated copies as separable uses that each required analysis. The settlement—about $3,000 per work across roughly 500,000 books—also shows how corpus-wide sourcing exposure can accumulate per work. Provenance therefore has to be resolved at acquisition, not inferred from a later training purpose.

09 — Prevention checklist

  • Gate corpus admission on verified lawful-acquisition status for every external dataset.
  • Record each source, acquisition date, license, and lawful basis in a provenance ledger.
  • Require legal review before bulk external data is copied into a standing corpus.
  • Give unverified sources a quarantine path and improperly sourced data a tested purge path.
  • Set a justified retention period for source copies instead of keeping them for optional use.

References

  1. Court record

  2. Docs

  3. News

  4. Blog

  5. Blog

  6. News

10 — Solutions (3)

Fixed this yourself? Add the steps that helped with Claude's training pipeline sourced millions of books from pirate sites.