Lawsuit filed against Adobe for using pirated books in AI training
Adobe is currently facing a legal battle over allegations that it used pirated books to train its SlimLM model, with author Elizabeth Lyon at the forefront of the lawsuit. The SlimLM model, optimized for document-related tasks on mobile devices, was trained using the SlimPajama-627B dataset, derived from the RedPajama dataset, which includes the Books3 dataset. The Books3 dataset, containing over 190,000 books, has been central to copyright lawsuits involving other tech giants like Apple and Salesforce. Lyon claims that Adobe’s use of SlimPajama, which includes Books3 content, constitutes training an AI model with copyrighted material without authorization, potentially infringing on authors’ rights.
This lawsuit sheds light on a larger issue prevalent in the development of AI systems, as companies strive to enhance the intelligence of their tools by utilizing vast amounts of data. Various companies, including Apple, Salesforce, and Anthropic, have previously faced legal challenges regarding the source and licensing of their training data. For instance, Anthropic recently settled a lawsuit by agreeing to pay US$1.5 billion over allegations of using pirated works to train their chatbot, Claude. The essential consideration in these legal proceedings is the necessity for extensive data to train AI models, leading companies to sometimes source data without ensuring the licensing or intellectual property rights of the content.
The implications of such legal disputes extend beyond the courtroom, posing real challenges for marketers and content creators who rely on AI tools for campaign acceleration. Marketers utilizing AI for various tasks such as generating blog content, automating customer support, or creating social media visuals must now reassess their practices in light of potential legal and ethical risks associated with the data powering these tools. To navigate this evolving landscape, marketing professionals should take the following steps:
1. Understand the Data Sources: Marketers should inquire about the origin and licensing of the data used to train AI models. Transparency regarding data sources is crucial in ensuring compliance and minimizing legal risks.
2. Assess AI Integration: Conduct an audit of AI-driven content workflows within your organization to determine which tools are in use and the types of content they generate. This documentation can help respond effectively to any legal challenges related to the origin of content.
3. Establish Legal Safeguards: Contracts with AI vendors should include indemnity clauses safeguarding companies from liability in cases where AI models were trained on infringing data. Proactive legal measures can protect businesses from legal entanglements in the future.
4. Implement Responsible AI Policies: Develop internal guidelines outlining the responsible use of AI within marketing and content teams. These policies should address transparency, attribution practices, and instances where human review is necessary to uphold compliance and integrity standards.
As AI becomes an integral component of marketing strategies, the imperative for marketers to ensure ethical and lawful use of these technologies is paramount. By staying informed, implementing strong policies, and being vigilant about data sources, marketers can mitigate risks and safeguard against potential legal challenges stemming from AI data misuse. The evolving landscape of AI development necessitates a proactive approach to compliance, transparency, and ethical practice, ensuring that marketers navigate this terrain responsibly and ethically.