Anthropic resolves lawsuit in US over use of pirated books for AI training

Anthropic has recently settled a class-action lawsuit in the United States related to the usage of pirated books for training artificial intelligence (AI). The issue surrounds the access to extensive data required for training large language models (LLMs) in the AI industry. Companies like Anthropic often argue that for the effective training of their models, they need access to vast databases, including datasets that may not be legally available, such as pirated books. This raises concerns about the ethical implications and the need for consent and fair compensation in the AI model training process.

The foundation of the Anthropic case lies in the ethical and legal frameworks that govern the use of copyrighted materials for AI training. MediaNama’s analysis underscores the significance of establishing clear guidelines that uphold intellectual property rights while fostering innovation in AI technology. By ensuring that AI companies obtain appropriate consent and fair compensation for using copyrighted materials, the industry can build trust with creators, authors, artists, publishers, and researchers who play a crucial role in shaping these technologies.

The recent development in the legal battle with Anthropic involves a settlement in the class-action lawsuit filed against the company by a group of authors. The case, known as Bartz vs. Anthropic, is scheduled for finalization in early September, following a ruling by a U.S. District Court in Northern California that touched on the fair use of copyrighted works in AI model training. In its judgment, the court acknowledged the fair use of purchased copyrighted books for training purposes but also ordered a trial to address the illegal use of pirated copies.

The lawsuit against Anthropic was initiated by authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, who accused the company of unlawfully utilizing their work to train LLMs. Amid the legal proceedings, the court clarified that Anthropic had engaged in piracy by downloading copyrighted books from sources like Library Genesis (LibGen) and Pirate Library Mirror. Anthropic altered its practices in 2024 after realizing the implications of using copyrighted materials for research, opting to purchase books from official retailers and distributors and refine the content for training purposes through tokenisation, a process used in natural language processing.

Key considerations from the legal perspective revolve around the provisions outlined in Section 107 of the U.S. Copyright Act, which dictates the criteria for evaluating fair use exemptions in copyright law. Factors such as the purpose of use, the nature of the copyrighted work, the amount used, and the impact on the market value of the work are carefully weighed to determine compliance with fair use principles. The court emphasized that while transforming physical copies of books into digital formats does constitute a transformative process, companies must adhere to ethical practices and respect intellectual property rights during AI model training.

Moving forward, the settlement process aims to address the concerns raised by the authors regarding the illegal use of their work in AI training. The Class Counsel is set to release a list of unlawfully downloaded works, allowing affected authors to participate in the lawsuit against Anthropic. Eligible authors may receive compensation based on specific criteria, including copyright ownership and registration. The potential outcomes of the trial could lead to monetary awards for the authors, in line with civil and criminal penalties under federal copyright laws.

In conclusion, the Anthropic case underscores the importance of ethical considerations in AI model training and the necessity of establishing boundaries to protect intellectual property rights. By adhering to legal frameworks, AI companies can continue to innovate while respecting the contributions of creators and authors in shaping the future of technology.