Adobe being sued in US for allegedly using pirated books in AI training

Elizabeth Lyon filed a lawsuit against Adobe Inc. in a California court on December 16, 2025, alleging that the company utilized unauthorized copies of her and other authors’ books to train its SlimLM AI models without consent or compensation.

The legal action, on behalf of affected US copyright holders, accuses Adobe of infringing copyrights of literary works to develop its small language model (SLM) for document-related tasks on mobile devices.

This lawsuit marks Adobe’s primary copyright dispute regarding AI training data, amidst increasing legal scrutiny on how generative AI systems access and use copyrighted material.

Earlier this year, Anthropic settled a class action with authors for $1.5 billion, claiming the company used pirated books to train large language models (LLMs), setting a record for copyright recovery. Companies like Apple, OpenAI, and Meta have also encountered allegations linked to their AI training methods.

The lawsuit claims Adobe violated authors’ copyrights by incorporating SlimLM series models in its SLMs, specifically designed for use on devices like smartphones, tablets, and laptops. These models handle document-assistance tasks as part of Adobe’s AI products.

The focal point of the dispute lies in the SlimPajama training dataset used by Adobe to pre-train SlimLM models. Plaintiffs argue that SlimPajama is a revised version of the RedPajama dataset, originating from the RedPajama-Books section of the Books3 dataset obtained from the Bibliotik private tracker.

Allegedly, Adobe retained these datasets on its servers, merging the extracted data into SlimLM models’ parameters even after the initial training, thus perpetuating copyright infringements.

The Books3 dataset, comprising nearly 200,000 pirated e-books from sources like Bibliotik, has been pivotal in AI copyright litigation in 2025. It features heavily as a legal issue, exemplified in a proposed class action against Apple over the use of Books3 in training its OpenELM model for Apple Intelligence.

In previous cases, Books3 material prompted a significant recovery against Anthropic and claims against Meta by Entrepreneur Media over the LLaMA models. These lawsuits challenge Meta’s usage of unauthorized materials like Books3 and Library Genesis (LibGen).

This legal battle highlights concerns about the construction and commercialization of AI systems that rely on third-party text corpora. The focus on Adobe’s SlimLM models addresses features integrated into mainstream productivity software, widening the implications beyond experimental tools.

The lawsuit emphasizes risks associated with derivative datasets such as SlimPajama, which may claim to be cleaned or deduplicated but still contain copyrighted content. Should the plaintiffs’ arguments prevail, AI companies may face exposure not just for direct infringement but for incorporating datasets originating from sources with copyright violations.