Some of the most infamous so-called shadow libraries, such as Z-Library and Library Genesis (Libgen), have faced increasing legal pressure to cease operations or risk being pushed to the dark web. Among their key critics are the US Department of Justice, which has charged Z-Library with criminal copyright infringement, and textbook publishers who sued Libgen last year for allegedly distributing copyrighted works on a massive scale in violation of copyright laws.
However, Nvidia has emerged as an unexpected defender of these shadow libraries in a recent lawsuit filed by book authors. The lawsuit concerns data repositories scraped to build the Books3 dataset, used to train Nvidia’s AI platform NeMo. Authors allege that Nvidia used data from notorious shadow libraries like Bibliotik, Z-Library, Libgen, Sci-Hub, and Anna’s Archive.
Nvidia seeks to invalidate the authors’ copyright claims by denying that these websites should even be classified as shadow libraries. “Nvidia denies the characterization of the listed data repositories as ‘shadow libraries’ and denies that hosting data in or distributing data from the repositories necessarily violates the US Copyright Act,” Nvidia stated in its court filing.
The chipmaker also refuted claims of improper use or copying, defending its AI training methods as fair use. Nvidia argued that “training is a highly transformative process” that involves adjusting numerical parameters, such as weights, and that the outputs of a large language model (LLM) are based on these weights.
Authors contend that these weights are derived from copyrighted material in the training dataset, used without their consent or compensation. Some companies, like OpenAI, have already started licensing publishers’ content to avoid these issues. Lawyers for The New York Times, currently suing OpenAI, suggest that OpenAI’s recent content licensing deal with News Corp demonstrates that publishers should be paid when their work is used for AI.
Until courts or lawmakers settle this debate, companies training AI with the Books3 dataset will likely continue to face lawsuits from rights holders. A lawyer for textbook publishers suing Libgen labeled the site a “thieves’ den” of illegal books, arguing that its conduct is “massively illegal.”
Nvidia’s stance on whether these controversial websites qualify as shadow libraries may not resonate with those running the sites. Anna, the pseudonymous creator of Anna’s Archive, openly describes the site as “the world’s largest shadow library.”
As shadow libraries advocate for the free distribution of information, AI companies like Nvidia might find themselves aligned with this ethos to maximize profits and dominate the AI market. Nvidia recently announced a record $26 billion revenue for the first quarter of 2024, underscoring the financial stakes involved.
For more details, visit the original article on Ars Technica.