Anthropic also acquired infringing copies of works from pirate sites. Judge Alsup ruled that these, and uses made from them, are not fair use.
A federal judge has issued a landmark fair use decision in a generative-AI copyright infringement lawsuit.
In a previous blog post, I wrote about the fair use decision in Thomson Reuters v. ROSS. As I explained there, that case involved a search-and-retrieval AI system, so the holding was not determinative of fair use in the context of generative AI. Now we finally have a decision that addresses fair use in the generative-AI context.
Bartz et al. v. Anthropic PBC
I did not include this case in my list of the top 12 generative-AI lawsuits, but only because it was one among many raising the same basic questions about training AI on copyright-protected works. This issue was well represented by others on the list. As it happens, though, Bartz has now taken on enhanced significance because the judge in the case has issued an important ruling on fair use.
Anthropic is an AI software firm founded by former OpenAI employees. It offers a generative-AI tool called Claude. Like other generative-AI tools, Claude mimics human conversational skills. When a user enters a text prompt, Claude will generate a response that is very much like one a human being might make (except it is sometimes more knowledgeable.) It is able to do this by using large language models (LLMs) that have been trained on millions of books and texts.
Adrea Bartz, Charles Graeber, and Kirk Wallace Johnson are book authors. In August 2024, they sued Anthropic, claiming the company infringed the copyrights in their works. Specifically, they alleged that Anthropic copied their works from pirated and purchased sources, digitized print versions, assembled them into a central library, and used the library to train LLMs, all without permission. Anthropic asserted, among other things, a fair use defense.
Earlier this year, Anthropic filed a motion for summary judgment on the question of fair use.
On June 23, 2025, Judge Alsup issued an order granting summary judgment in part and denying it in part. It is the first major ruling on fair use in the dozens of generative-AI copyright infringement lawsuits that are currently pending in federal courts.
The Order includes several key rulings.
Digitization
Anthropic acquired both pirated and lawfully purchased printed copies of copyright-protected works and digitized them to create a central e-library. Authors claimed that making digital copies of their works infringed the exclusive right of copyright owners to reproduce their works. (See 17 U.S.C. 106.)
In the process of scanning print books to create digital versions of them, the print copies were destroyed. Book bindings were stripped so that each individual page could be scanned. The print copies were then discarded. The digital copies were not distributed to others. Under these circumstances, the court ruled that making digital versions of print books is fair use.
The court likened format to a frame around a work, as distinguished from the work itself. As such, a digital version is not a new derivative work. Rather, it is a transformative use of an existing work. So long as the digital version is merely a substitute for a print version a person has lawfully acquired, and so long as the print version is destroyed and the digital version is not further copied or distributed to others, then digitizing a printed work is fair use. This is consistent with the first sale doctrine (17 U.S.C. 109(a)), which gives the purchaser of a copy of a work a right to dispose of that particular copy as the purchaser sees fit.
In short, the mere conversion of a lawfully acquired print book to a digital file to save space and enable searchability is transformative, and so long as the print version is destroyed and the digital version is not further copied or distributed, it is fair use.
AI Training Is Transformative Fair Use
The authors did not contend that Claude generated infringing output. Instead, they argued that copies of their works were used as inputs to train the AI. The Copyright Act, however, does not prohibit or restrict the reading or analysis of copyrighted works. So long as a copy is lawfully purchased, the owner of the purchased copy can read it and think about it as often as he or she wishes.
[I]f someone were to read all the modern-day classics because of their exceptional expression, memorize them, and then emulate a blend of their best writing, would that violate the Copyright Act? Of course not.
Judge Alsup described AI training as “spectacularly” transformative.” Id. After considering all four fair use factors, he concluded that training AI on lawfully acquired copyright-protected works (as distinguished from the initial acquisition of copies) is fair use.
Pirating Is Not Fair Use
In addition to lawfully purchasing copies of some works, Anthropic also acquired infringing copies of works from pirate sites. Judge Alsup ruled that these, and uses made from them, are not fair use. The case will now proceed to trial on the issue of damages resulting from the infringement.
Conclusion
Each of these rulings seems, well, sort of obvious. It is nice to have the explanations laid out so clearly in one place, though.
A court has handed down the first known ruling (to me, anyway) on “fair use” in the wave of copyright infringement lawsuits against AI companies that are pending in federal courts.
Thomson Reuters v. ROSS is one of the top 12 generative-AI lawsuits that are pending in the courts. A court has handed down the first known ruling (to me, anyway) on “fair use” in the wave of copyright infringement lawsuits against AI companies that are pending in federal courts. The ruling came in Thomson Reuters v. ROSS. Thomas Reuters filed this lawsuit against Ross Intelligence back in 2020, alleging that Ross trained its AI models on Westlaw headnotes to build a competing legal research tool, infringing numerous copyrights in the process. Ross asserted a fair use defense.
Library of Congress
In 2023, Thomson Reuters sought summary judgment against Ross on the fair use defense. At that time, Judge Bibas denied the motion. This week, however, the judge reversed himself, knocking out at least a major portion of the fair use defense.
Ross had argued that Westlaw headnotes are not sufficiently original to warrant copyright protection and that even if they are, the use made of them was “fair use.” After painstakingly reviewing the headnotes and comparing them with the database materials, he concluded that 2,243 headnotes were sufficiently original to receive copyright protection, that Ross infringed them, and that “fair use” was not a defense in this instance because the purpose of the use was commercial and it competed in the same market with Westlaw. Because of that, it was likely to have an adverse impact on the market for Westlaw.
While this might seem to spell the end for AI companies in the many other lawsuits where they are relying on a “fair use” defense, that is not necessarily so. As Judge Bibas noted, the Ross AI was non-generative. Generative AI tools may be distinguishable in the fair use analysis.
I will be presenting a program on Recent Developments in AI Law in New Jersey this summer. This one certainly will merit mention. Whether any more major developments will come to pass between now and then remains to be seen.
New AI Copyright Infringement Lawsuit
Another copyright and trademark infringement lawsuit against an AI company was filed this week. Advance Local Media et al. v. Cohere, Inc. This one pits news article publishers Advance Local Media, Condé Nast, The Atlantic, Forbes Media, The Guardian, Business Insider, LA Times, McClatchy Media Company, Newsday, Plain Dealer Publishing Company, POLITICO, The Republican Company, Toronto Star Newspapers, and Vox Media against AI company Cohere.
The complaint alleges that Cohere made unauthorized use of publisher content in developing and operating its generative AI systems, infringing numerous copyrights and trademarks. The plaintiffs are seeking an injunction and monetary damages.
A status update on 24 pending lawsuits against AI companies – what they’re about and what is happening in court – prepared by Minnesota copyright attorney Thomas James.
Advancements in artificial intelligence technology, including generative-AI, have introduced a wide range of new or exacerbated legal problems. Collectively, I call these AI legal issues. Although not all of them are unique to scenarios involving AI, they are certainly testing and stretching the capacity of legal institutions. Here is a very brief summary of how these issues are playing out in the courts, as of February 28, 2024.
Copyright
Thaler v. Perlmutter (D.D.C. 2022).
Complaint filed June 2, 2022. Thaler created an AI system called the Creativity Machine. He applied to register copyrights in the output he generated with it. The Copyright Office refused registration on the ground that AI output does not meet the “human authorship” requirement. (I explained that requirement in a previous blog post that explored the difference between human and AI creation of a work. He then sought judicial review. The district court granted summary judgment for the Copyright Office. (SeeA Recent Exit from Paradise.) In October, 2023, Thaler filed an appeal to the District of Columbia Circuit Court of Appeals (Case no. 23-5233).
Doe v. GitHub, Microsoft, and OpenAI (N.D. Cal. 2022)
I wrote about this case in Generative AI: The Top 12 Lawsuits. The complaint was filed November 3, 2022. Software developers claim the defendants trained Codex and Copilot on code derived from theirs, which they published on GitHub. Some claims have been dismissed, but claims that GitHub and OpenAI violated the DMCA and breached open source licenses remain. Discovery is ongoing.
Andersen v. Stability AI (N.D. Cal. 2023)
The Andersen v. Stability AI complaint was filed January 13, 1023. Visual artists sued Midjourney, Stability AI and DeviantArt for copyright infringement for allegedly training their generative-AI models on images scraped from the Internet without copyright holders’ permission. Other claims included DMCA violations, publicity rights violations, unfair competition, breach of contract, and a claim that output images are infringing derivative works. On October 30, 2023, the court largely granted motions to dismiss, but granted leave to amend the complaint. Plaintiffs filed an amended complaint on November 29, 2023. Defendants have filed motions to dismiss the amended complaint. Hearing on the motion is set for May 8, 2024.
Getty Images v. StabilityAI (U.K. 2023)
Complaint filed January, 2023. Getty Images claims StabilityAI scraped images without its consent. The Getty Images lawsuit has survived a motion to dismiss. The case appears to be heading to trial.
In re OpenAI ChatGPT Litigation (N.D. Cal. 2023)
Complaint filed June 28, 3023. Originally captioned Tremblay v. OpenAI. Book authors sued OpenAI for direct and vicarious copyright infringement, DMCA violations, unfair competition and negligence. Both input (training) and output (derivative works) claims are alleged, as well as state law claims of unfair competition, etc.
Most state law and DMCA claims have been dismissed, but claims based on unauthorized copying during the AI training process remain. An amended complaint is likely to come in March. The court has directed the amended complaint to consolidate Tremblay v. OpenAI, Chabon v. OpenAI, and Silverman v. OpenAI.
Kadrey v. Meta (N.D. Cal. 2023)
Complaint filed July 7, 2023. Sarah Silverman and other authors allege Meta infringed copyrights in their works by making copies of them while training Meta’s AI model; that the AI model is itself an infringing derivative work; and that outputs are infringing copies of their works. Plaintiffs also allege DMCA violations, unfair competition, unjust enrichment, and negligence. The court granted Meta’s motion to dismiss all claims except the claim that unauthorized copies were made during the AI training process. An amended complaint and answer have been filed.
In 2025, Judge Chhabria ruled in Meta’s favor on fair use with respect to AI training; reserved the motion for summary judgment on the DMCA claims for decision in a separate order, and held that the claim of infringing distribution via leeching or seeding “will remain a live issue in the case.” Kadrey v. Meta Platforms.
J.L. v. Google (N.D. Cal. 2023)
Complaint filed July 11, 2023. In another case I mentioned in Generative AI: The Top 12 Lawsuits, an author filed a complaint against Google alleging misuse of content posted on social media and Google platforms to train Google’s AI Bard. (Gemini is the successor to Google’s Bard.) Claims include copyright infringement, DMCA violations, and others. J.L. filed an amended complaint and Google has filed a motion to dismiss it. A hearing is scheduled for May 16, 2024.
Chabon v. OpenAI (N.D. Cal. 2023)
Complaint filed September 9, 2023. Authors allege that OpenAI infringed copyrights while training ChatGPT, and that ChatGPT is itself an unauthorized derivative work. They also assert claims of DMCA violations, unfair competition, negligence and unjust enrichment. Chabon v. OpenAI has been consolidated with Tremblay v. OpenAI, and the cases are now captioned In re OpenAI ChatGPT Litigation.
Chabon v. Meta Platforms (N.D. Cal. 2023)
Complaint filed September 12, 2023. Authors assert copyright infringement claims against Meta, alleging that Meta trained its AI using their works and that the AI model itself is an unauthorized derivative work. The authors also assert claims for DMCA violations, unfair competition, negligence, and unjust enrichment. In November, 2023, the court issued an Order dismissing all claims except the claim of unauthorized copying in the course of training the AI. The court described the claim that an AI model trained on a work is a derivative of that work as “nonsensical.” Chabon v. Mea Platforms.
Authors Guild v. OpenAI, Microsoft, et al. (S.D.N.Y. 2023)
Complaint filed September 19, 1023. Book and fiction writers filed a complaint for copyright infringement in connection with defendants’ training AI on copies of their works without permission. A motion to dismiss has been filed. Authors Guild v. Open AI et al.
Huckabee v. Bloomberg, Meta Platforms, Microsoft, and EleutherAI Institute (S.D.N.Y. 2023)
Complaint filed October 17, 2023. Political figure Mike Huckabee and others allege that the defendants trained AI tools on their works without permission when they used Books3, a text dataset compiled by developers; that their tools are themselves unauthorized derivative works; and that every output of their tools is an infringing derivative work. Claims against EleutherAI have been voluntarily dismissed. Claims against Meta and Microsoft have been transferred to the Northern District of California. Bloomberg is expected to file a motion to dismiss soon. Huckabee v. Bloomberg et al.
Huckabee v. Meta Platforms and Microsoft (N.D. Cal. 2023)
Complaint filed October 17, 2023. Political figure Mike Huckabee and others allege that the defendants trained AI tools on their works without permission when they used Books3, a text dataset compiled by developers; that their tools are themselves unauthorized derivative works; and that every output of their tools is an infringing derivative work. Plaintiffs have filed an amended complaint. Plaintiffs have stipulated to dismissal of claims against Microsoft without prejudice. Huckabee v. Meta Platforms and Microsoft.
Concord Music Group v. Anthropic (M.D. Tenn. 2023)
Complaint filed October 18, 2023. Music publishers claim that Anthropic infringed publisher-owned copyrights in song lyrics when they allegedly were copied as part of an AI training process (Claude) and when lyrics were reproduced and distributed in response to prompts. They have also made claims of contributory and vicarious infringement. Motions to dismiss and for a preliminary injunction are pending. Concord Music Group v. Anthropic.
Alter v. OpenAI and Microsoft (S.D.N.Y. 2023)
Complaint filed November 21, 2023. Nonfiction author alleges claims of copyright infringement and contributory copyright infringement against OpenAI and Microsoft, alleging that reproducing copies of their works in datasets used to train AI infringed copyrights. The court has ordered consolidation of Author’s Guild (23-cv-8292) and Alter (23-cv-10211). On February 12,2024, plaintiffs in other cases filed a motion to intervene and dismiss. Alter v. OpenAI and Microsoft.
New York Times v. Microsoft and OpenAI (S.D.N.Y. 2023)
Complaint filed December 27, 2023. The New York Times alleges that their news stories were used to train AI without a license or permission, in violation of their exclusive rights of reproduction and public display, as copyright owners. The complaint also alleges vicarious and contributory copyright infringement, DMCA violations, unfair competition, and trademark dilution. The Times seeks damages, an injunction against further infringing conduct, and a Section 503(b) order for the destruction of “all GPT or other LLM models and training sets that incorporate Times Works.” On February 23, 2024, plaintiffs in other cases filed a motion to intervene and dismiss this case. New York Times v. Microsoft and OpenAI.
Basbanes and Ngagoyeanes v. Microsoft and OpenAI (S.D.N.Y. 2024)
Getty Images v. Stability AI Complaint filed on February 3, 2023. Getty Images alleges claims of copyright infringement, DMCA violation and trademark violations against Stability AI. The judge has dismissed without prejudice a motion to dismiss or transfer on jurisdictional grounds. The motion may be re-filed after the conclusion of jurisdictional discovery, which is ongoing.
Privacy and Publicity Rights
Flora v. Prisma Labs (N.D. Cal.)
Complaint filed February 15, 2023. Plaintiffs allege violations of the Illinois Biometric Privacy Act in connection with Prisma Labs’ collection and retention of users’ selfies in AI training. The court has granted Prisma’s motion to compel arbitration. Flora v. Prisma Labs.
Kyland Young v. NeoCortext (C.D. Cal. 2023)
Complaint filed April 3, 2023. This complaint alleges that AI tool Reface used a person’s image without consent, in violation of the person’s publicity rights under California law. The court has denied a motion to dismiss, ruling that publicity rights claims are not preempted by federal copyright law. The case has been stayed pending appeal. Kyland Young v. NeoCortext.
Complaint filed June 28, 2023. Users claim OpenAI violated the federal Electronic Communications Privacy Act and California wiretapping laws by collecting their data when they input content into ChatGPT. They also claim violations of the Computer Fraud and Abuse Act. Plaintiffs voluntarily dismissed the case on September 15, 2023. See now A.T. v. OpenAI (N.D. Cal. 2023) (below). P.M. v. OpenAI.
A.T. v. OpenAI (N.D. Cal. 2023)
Complaint filed September 5, 2023. ChatGPT users claim the company violated the federal Electronic Communications Privacy Act, the Computer Fraud and Abuse Act, and California Penal Code section 631 (wiretapping). The gravamen of the complaint is that ChatGPT allegedly accessed users’ platform access and intercepted their private information without their knowledge or consent. Motions to dismiss and to compel arbitration are pending. A.T. v. OpenAI.
Defamation
Walters v. OpenAI (Gwinnett County Super. Ct. 2023), and Walters v. OpenAI (N.D. Ga. 2023)
Gwinnett County complaint filed June 5, 2023.
Federal district court complaint filed July 14, 2023.
Radio Radio talk show host sued OpenAI for defamation. A reporter had used ChatGPT to get information about him. ChatGPT wrongly described him as a person who had been accused of fraud. In October, 2023, the federal court remanded the case to the Superior Court of Gwinnett County, Georgia. On January 11, 2024, the Gwinnett County Superior Court denied OpenAI’s motion to dismiss. Walters v. OpenAI.
Battle v. Microsoft (D. Md. 2023)
Complaint filed July 7, 2023. Pro se defamation complaint against Microsoft alleging that Bing falsely described him as a member of the “Portland Seven,” a group of Americans who tried to join the Taliban after 9/11. Battle v. Microsoft.
Caveat
This list is not exhaustive. There may be other cases involving AI that are not included here. For a discussion of bias issues in Google’s Gemini, have a look at Scraping Bias.
Contains information related to marketing campaigns of the user. These are shared with Google AdWords / Google Ads when the Google Ads and Google Analytics accounts are linked together.
90 days
__utma
ID used to identify users and sessions
2 years after last activity
__utmt
Used to monitor number of Google Analytics server requests
10 minutes
__utmb
Used to distinguish new sessions and visits. This cookie is set when the GA.js javascript library is loaded and there is no existing __utmb cookie. The cookie is updated every time data is sent to the Google Analytics server.
30 minutes after last activity
__utmc
Used only with old Urchin versions of Google Analytics and not with GA.js. Was used to distinguish between new sessions and visits at the end of a session.
End of session (browser)
__utmz
Contains information about the traffic source or campaign that directed user to the website. The cookie is set when the GA.js javascript is loaded and updated when data is sent to the Google Anaytics server
6 months after last activity
__utmv
Contains custom information set by the web developer via the _setCustomVar method in Google Analytics. This cookie is updated every time new data is sent to the Google Analytics server.
2 years after last activity
__utmx
Used to determine whether a user is included in an A / B or Multivariate test.
18 months
_ga
ID used to identify users
2 years
_gali
Used by Google Analytics to determine which links on a page are being clicked
30 seconds
_ga_
ID used to identify users
2 years
_gid
ID used to identify users for 24 hours after last activity
24 hours
_gat
Used to monitor number of Google Analytics server requests when using Google Tag Manager