WikiHow has initiated legal action against OpenAI, claiming that the company has been using its copyrighted how-to articles to train ChatGPT, leading to the generation of similar or nearly identical content. The case has been brought to a New York federal court, with WikiHow asserting that the model reproduces its distinctive instructional articles verbatim.
This lawsuit underscores ongoing concerns over how these AI systems are trained. The alleged mass-scale copying of content raises questions about intellectual property rights and the ethical use of data in AI development. OpenAI’s ChatGPT, a focal point in this legal challenge, is among the large language models often trained on vast datasets, which sometimes include publicly available internet content.
The tension between content creators and AI developers is not new. As the AI ecosystem develops, companies have increasingly faced scrutiny over their practices in building these advanced systems. This case follows a pattern where content producers have accused tech companies of exploiting their copyrighted work without permission or compensation, as seen in other legal disputes, such as those involving Getty Images and Stability AI.
Beyond the courtroom, the implications for AI training practices are profound. Should the court side with WikiHow, the decision may establish precedents for how copyrighted materials can be used in AI training. Legal experts are closely watching the case for indications of how intellectual property law might evolve to address these cutting-edge issues.
WikiHow’s assertions, detailed in the lawsuit, reflect broader industry concerns about balancing innovation with rights protection. As the legal landscape continues to adjust to these challenges, both AI developers and rights holders will need to navigate a complex set of emerging standards and practices.