À propos de la propriété intellectuelle Formation en propriété intellectuelle Respect de la propriété intellectuelle Sensibilisation à la propriété intellectuelle La propriété intellectuelle pour… Propriété intellectuelle et… Propriété intellectuelle et… Information relative aux brevets et à la technologie Information en matière de marques Information en matière de dessins et modèles Information en matière d’indications géographiques Information en matière de protection des obtentions végétales (UPOV) Lois, traités et jugements dans le domaine de la propriété intellectuelle Ressources relatives à la propriété intellectuelle Rapports sur la propriété intellectuelle Protection des brevets Protection des marques Protection des dessins et modèles Protection des indications géographiques Protection des obtentions végétales (UPOV) Règlement extrajudiciaire des litiges Solutions opérationnelles à l’intention des offices de propriété intellectuelle Paiement de services de propriété intellectuelle Décisions et négociations Coopération en matière de propriété intellectuelle Appui à l’innovation Partenariats public-privé L’Organisation L’OMPI et l’intelligence artificielle Travailler à l’OMPI Responsabilité Brevets Marques Dessins et modèles Indications géographiques Droit d’auteur Secrets d’affaires Avenir de la propriété intellectuelle Académie de l’OMPI Ateliers et séminaires Application des droits de propriété intellectuelle WIPO ALERT Sensibilisation Journée mondiale de la propriété intellectuelle Magazine de l’OMPI Études de cas et exemples de réussite Actualités dans le domaine de la propriété intellectuelle Prix de l’OMPI Entreprises Femmes Universités Peuples autochtones Instances judiciaires Jeunesse Examinateurs Écosystèmes d’innovation Économie Financement Actifs incorporels Santé mondiale Changement climatique Politique en matière de concurrence Objectifs de développement durable Ressources génétiques, savoirs traditionnels et expressions culturelles traditionnelles Technologies de pointe Applications mobiles Sport Tourisme Musique Mode PATENTSCOPE Analyse de brevets Classification internationale des brevets Programme ARDI – Recherche pour l’innovation Programme ASPI – Information spécialisée en matière de brevets Base de données mondiale sur les marques Madrid Monitor Base de données Article 6ter Express Classification de Nice Classification de Vienne Base de données mondiale sur les dessins et modèles Bulletin des dessins et modèles internationaux Base de données Hague Express Classification de Locarno Base de données Lisbon Express Base de données mondiale sur les marques relative aux indications géographiques Base de données PLUTO sur les variétés végétales Base de données GENIE Traités administrés par l’OMPI WIPO Lex – lois, traités et jugements en matière de propriété intellectuelle Normes de l’OMPI Statistiques de propriété intellectuelle WIPO Pearl (Terminologie) Publications de l’OMPI Profils nationaux Centre de connaissances de l’OMPI Données essentielles sur l’investissement incorporel dans le monde Série de rapports de l’OMPI consacrés aux tendances technologiques Indice mondial de l’innovation Rapport sur la propriété intellectuelle dans le monde PCT – Le système international des brevets ePCT Budapest – Le système international de dépôt des micro-organismes Madrid – Le système international des marques eMadrid Article 6ter (armoiries, drapeaux, emblèmes nationaux) La Haye – Le système international des dessins et modèles industriels eHague Lisbonne – Le système d’enregistrement international des indications géographiques eLisbon UPOV PRISMA Médiation Arbitrage Procédure d’expertise Litiges relatifs aux noms de domaine Accès centralisé aux résultats de la recherche et de l’examen (WIPO CASE) Service d’accès numérique aux documents de priorité (DAS) WIPO Pay WIPO Wallet Assemblées de l’OMPI Comités permanents Calendrier des réunions WIPO Webcast Documents officiels de l’OMPI Plan d’action de l’OMPI pour le développement Initiatives et projets sur mesure Forums de collaboration et dialogues Programme d’accélération pour l’innovation, la créativité et le développement La propriété intellectuelle en action Stratégies nationales de propriété intellectuelle et d’innovation Pôle de coopération Centres d’appui à la technologie et à l’innovation (CATI) Transfert de technologie Programme d’aide aux inventeurs Commercialisation de la propriété intellectuelle WIPO GREEN Initiative PAT-INFORMED de l’OMPI Consortium pour des livres accessibles L’OMPI pour les créateurs États membres Observateurs Directeur général Activités par unité administrative Bureaux extérieurs Forum mondial sur la propriété intellectuelle et l’intelligence artificielle Plateforme d’échange sur l’infrastructure de l’intelligence artificielle Outils et services en matière d’intelligence artificielle Postes de fonctionnaires Postes de personnel affilié Achats Résultats et budget Rapports financiers Audit et supervision
Arabic English Spanish French Russian Chinese
Lois Traités Jugements Recherche par ressort juridique

États-Unis d'Amérique

US161-j

Retour

2026 WIPO IP Judges Forum Informal Case Summary – United States District Court for the Northern District of California [2025]: Bartz v Anthropic PBC, 787 F.Supp.3d 1007

This is an informal case summary prepared for the purposes of facilitating exchange during the 2026 WIPO IP Judges Forum.

 

Session 3: Copyright and AI Model Training

 

United States District Court for the Northern District of California [2025]: Bartz v Anthropic PBC, 787 F.Supp.3d 1007 

 

Date of judgment: June 23, 2025

Issuing authority: United States District Court for the Northern District of California

Level of the issuing authority: First Instance

Type of procedure: Judicial (Civil)

Subject matter: Copyright and Related Rights (Neighboring Rights)

Plaintiff/Appellant: Andrea Bartz; Charles Graeber; Kirk Wallace Johnson; Affiliated corporate entities

Defendant/Respondent: Anthropic PBC

Keywords: Artificial intelligence (AI); Fair use; Transformative use; Large-language-model training; Pirated copies; Destructive scanning; Format shifting; Intermediate copying; Digital libraries

 

Basic facts: Anthropic is an artificial-intelligence company whose principal product is the Claude service, powered by large language models (LLMs). The models were trained on text selected from a central research library assembled by the company. Anthropic sought to create a permanent collection of what it described as “all the books in the world,” retained “forever.”

 

Anthropic acquired books for that library through two principal routes. First, between January 2021 and July 2022, it downloaded more than seven million digital copies of books from pirate sources, including Books3, LibGen, and PiLiMi. The court found that Anthropic knew that those sources supplied unauthorized copies. Some of the pirated books were later used for model training, but others were not; Anthropic retained the library copies regardless of whether a particular work had been selected for training or was expected to be used in future training.

 

Second, in spring 2024, Anthropic began purchasing millions of print books. It removed the bindings, cut the books apart for scanning, created digital copies containing page images and machine-readable text, and discarded the physical originals. Anthropic retained the resulting searchable digital files in its internal research library. These digital files were not distributed, shown, or sold outside the company.

 

From the central library, Anthropic selected collections and subsets of books for particular “data mixes.” These materials were cleaned, tokenized, and processed to train successive LLMs. For purposes of the motion, the court assumed the plaintiffs’ allegation that the resulting models retained compressed or “memorized” versions of the works that could, under some circumstances, reproduce them almost verbatim. The plaintiffs did not, however, allege that Anthropic had delivered infringing outputs to the public or that any output supplied to a user was traceable to one of their works.

 

The named authors – Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson – brought a putative copyright class action in August 2024. Before class certification, the court permitted Anthropic to move for summary judgment on the discrete issue of fair use.

 

Held: The court separated Anthropic’s conduct into three relevant uses:

1.    Copying works to train particular LLMs.

2.    Creating and retaining digital library copies by destructively scanning lawfully purchased print books.

3.    Acquiring and indefinitely retaining pirated digital copies in a permanent, general-purpose central library.

 

The court granted summary judgment for Anthropic on the first two categories. It held that the copying used to train LLMs was fair use, including any alleged memorization incident to the training process. It also held that Anthropic’s destructive scanning and internal retention of digital replacements for lawfully purchased print copies was fair use on the facts before it.

 

The court denied Anthropic’s motion for summary judgment regarding the pirated-library copies. It rejected Anthropic’s argument that those copies should be excused as part of, or preparatory to, otherwise fair-use model training. Building and indefinitely retaining a comprehensive library of pirated books for possible present or future use was a distinct use, and the court concluded that it was not fair use.

 

The order did not enter final judgment for the plaintiffs on the pirated-copy claims and did not resolve liability, damages, statutory damages, willfulness, class certification, or other remaining issues. The court directed that the pirated-copy claims proceed to trial on liability and on damages, actual or statutory, including willfulness. The fact that Anthropic later acquired authorized copies of books previously obtained from pirate sources did not retroactively cure the earlier infringement, although it could bear on remedies, including statutory damages and willfulness.

 

Relevant holdings in relation to copyright and AI training: Applying the four statutory factors in section 107 of the Copyright Act, and following Andy Warhol Foundation for the Visual Arts, Inc v Goldsmith, the court accepted the plaintiffs’ insistence that Anthropic’s copying involved distinct uses. The relevant distinction was between copying to build and retain a central library of books and copying selected works or subsets of works to train a particular LLM. The plaintiffs separately challenged the conversion of lawfully acquired print books into replacement digital copies.

 

Training copies

 

For the LLM-training use, the first factor—the purpose and character of the use—strongly favored fair use. The court described the use as “transformative—spectacularly so” and as “quintessentially transformative.” Anthropic used the works to identify and model statistical relationships among textual fragments, enabling the completed model to receive new prompts and generate new responses. The court regarded that function as fundamentally different from the ordinary expressive purpose of the books.

 

The court rejected the proposition that copyright owners may prevent others from using works as material from which to learn or to produce new works. It reasoned that copyright does not give an author control over the ideas, principles, concepts, or methods expressed in a work, and that a rule requiring payment whenever a person recalled a work or drew upon it in new writing would be untenable.

 

The court distinguished Thomson Reuters Enterprise Centre GmbH v Ross Intelligence Inc, emphasizing that the system at issue there was not generative AI and was designed to provide legal research results, including court opinions, in response to legal queries. By contrast, the court treated an AI system trained to generate fresh text as analogous to a person who has learned from books and later produces new writing. It also considered Anthropic’s output filters relevant to the first factor because they supported the absence of allegedly infringing output delivered to users; it analogized that constraint to the snippet restrictions that prevented Google Books from becoming a substitute reading service.

The second factor – the nature of the copyrighted works – favored the authors. The works were expressive books, and Anthropic had selected books precisely for their expressive content. The court regarded this factor as having limited independent weight, principally because it informed the other factors.

 

The third factor—the amount and substantiality of the portion used—favored fair use for training. Complete copying was reasonably necessary to the transformative purpose of using very large bodies of text to train an LLM. “Reasonably necessary” did not mean strictly indispensable, and the court rejected the suggestion that Anthropic had to demonstrate why each individual work, rather than some alternative work, was needed. The court also stressed that the plaintiffs had not alleged that complete copies, or substantial portions, were made available to the public through infringing outputs.

 

On the fourth factor—the effect on the potential market for or value of the copyrighted works—the court assumed for purposes of the motion that LLMs could facilitate a substantial proliferation of competing works and that a market for licensing books as AI-training data might emerge. It nevertheless held that these propositions did not establish cognizable market harm. Copyright protects the market for an author’s expression and against substitution for the author’s work; it does not protect authors from competition created by new works produced by persons or systems that have learned from existing works. The court analogized the claimed harm to the non-actionable effect of teaching schoolchildren to write better. It also held that a prospective licensing market for training uses was not a market that copyright law entitled rightsholders to reserve.

 

Balancing the factors, the court granted Anthropic summary judgment that copying works to train its LLMs was fair use.

 

Purchased and scanned books

 

The court treated the destructive scanning of lawfully purchased print books as a narrower, separate use. It rejected the argument that creation of the library was simply inseparable from model training and instead analyzed the print-to-digital conversion on its own terms.

Anthropic had lawfully purchased the physical books, destroyed the print originals while scanning them, and retained internal searchable digital replacements. The court reasoned that Anthropic was entitled, under the first-sale principle in section 109(a), to retain or dispose of the purchased physical copies. The scanning created a replacement digital copy rather than an additional copy retained alongside the original; no new creative content was added, no derivative work was created, and the digital files were not distributed outside Anthropic.

 

The court held that this form of format shifting was fair use. In particular, it regarded the digital replacement as transformative in the limited sense that it made the acquired books searchable and allowed Anthropic to reduce physical-storage burdens, while leaving the work’s expressive content unchanged and making no copy publicly available. The second factor favored the authors because the books were expressive, but the third factor favored fair use, because complete copying was necessary to make a functional searchable replacement, and the fourth factor was neutral, because the internal replacement files, each of which replaced a print copy Anthropic had purchased, did not displace any sale of the books.

 

 

 

 

Pirated central-library copies

 

The pirated-library copies stood differently. Anthropic had acquired millions of unauthorized digital books to assemble a permanent, general-purpose collection of material that it could retain and draw upon in present or future projects. The court held that this library-building and retention function was a use distinct from the specific use of selected books to train a particular LLM.

 

The court concluded that the pirated library was not transformative. Its purpose was to supply Anthropic with permanent access to complete works that it would otherwise have needed to acquire lawfully. Not every pirated book was used in model training; no general deletion practice removed books that had not been used or would never be used; and the library was maintained as a continuing resource for other or future uses.

 

The court concluded that all four factors weighed against fair use for this category. The first factor weighed against Anthropic because building and retaining a general-purpose library of pirated works was not transformative. The second factor favored the authors because the works were expressive. The third factor weighed against Anthropic because it had taken complete works. The fourth factor weighed against Anthropic because permitting such piracy would impair the ordinary market in which users acquire authorized copies of books.

 

Accordingly, the court denied Anthropic summary judgment on the pirated-library claims. It did not, in that order, enter summary judgment for the plaintiffs or determine damages or other remedies.

                                                                                      

Relevant legislation: United States Code, Title 17 – Copyrights (US455); Constitution of the United States of America (US431)