{"id":75896,"date":"2024-03-12T09:00:35","date_gmt":"2024-03-12T09:00:35","guid":{"rendered":"https:\/\/www.cryptocabaret.com\/?p=75896"},"modified":"2024-03-12T09:00:35","modified_gmt":"2024-03-12T09:00:35","slug":"authors-sue-nvidia-for-training-ai-on-pirated-books","status":"publish","type":"post","link":"https:\/\/www.cryptocabaret.com\/?p=75896","title":{"rendered":"Authors Sue NVIDIA for Training AI on Pirated Books"},"content":{"rendered":"<p><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2024\/03\/nvidia-logo-300x197.jpg\" alt=\"nvidia logo\" width=\"300\" height=\"197\" class=\"alignright size-medium wp-image-248322\" srcset=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2024\/03\/nvidia-logo-300x197.jpg 300w, https:\/\/torrentfreak.com\/images\/nvidia-logo.jpg 665w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\">Starting last year, various rightsholders have filed <a href=\"https:\/\/torrentfreak.com\/tag\/ai\/\">lawsuits<\/a> against companies that develop AI models.<\/p>\n<p>The list of complainants includes record labels, book authors, visual artists, even the New York Times. These rightsholders all object to the presumed use of their work without proper compensation.<\/p>\n<h2>\u201cBooks3\u201d<\/h2>\n<p>Many of the lawsuits filed by book authors come with a clear piracy angle. The cases allege that tech companies, including Meta, Microsoft, and OpenAI, used the controversial <a href=\"https:\/\/torrentfreak.com\/books3-takedown-anti-piracy-group-calls-for-more-ai-training-transparency-230905\/\">\u2018Books3\u2019 dataset<\/a> to train their models.<\/p>\n<p>Books3 was created by AI researcher Shawn Presser in 2020, who scraped the library of \u2018pirate\u2019 site Bibliotik. The dataset was broadly shared online and added to other databases including \u2018The Pile\u2018, an AI training dataset compiled by EleutherAI.<\/p>\n<p>After <a href=\"https:\/\/torrentfreak.com\/anti-piracy-group-takes-prominent-ai-training-dataset-books3-offline-230816\/\">pushback<\/a> from rightsholders and anti-piracy outfits, Books3 was taken offline over copyright concerns. However, for many of the companies that allegedly trained their AI models on it, there are still some legal repercussions to sort out. <\/p>\n<h2>Authors Sue NVIDIA for Copyright Infringement<\/h2>\n<p>On Friday, American authors <a href=\"https:\/\/en.wikipedia.org\/wiki\/Abdi_Nazemian\">Abdi Nazemian<\/a>, <a href=\"https:\/\/en.wikipedia.org\/wiki\/Brian_Keene\">Brian Keene<\/a>, and <a href=\"https:\/\/en.wikipedia.org\/wiki\/Stewart_O%27Nan\">Stewart O\u2019Nan<\/a> joined the barrage of legal action with a copyright infringement lawsuit against NVIDIA. The company, whose market cap exceeds $2 trillion, is mostly known for its GPUs and related software and services, but also has its own AI models. <\/p>\n<p>In a concise class action complaint, filed at a California federal court, the authors allege that NVIDIA used the Books3 dataset to train its <a href=\"https:\/\/docs.nvidia.com\/deeplearning\/nemo\/user-guide\/docs\/en\/main\/nlp\/megatron.html\">NeMo Megatron<\/a> language models. The models are hosted on Hugging Face where it <a href=\"https:\/\/huggingface.co\/nvidia\/nemo-megatron-gpt-1.3B#training-data\">states<\/a> that they are trained on EleutherAI\u2019s \u2018The Pile\u2019 dataset, which includes the pirated books.<\/p>\n<\/p>\n<p><center><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2024\/03\/nvidia-cvlaims.jpg\" alt=\"nvidia\" width=\"600\" height=\"443\" class=\"alignnone size-full wp-image-248326\" srcset=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2024\/03\/nvidia-cvlaims.jpg 1173w, https:\/\/torrentfreak.com\/images\/nvidia-cvlaims-300x221.jpg 300w\" sizes=\"auto, (max-width: 600px) 100vw, 600px\"><\/center><\/p>\n<p>Putting two and two together, the plaintiffs conclude that NVIDIA\u2019s models were trained on pirated books, including theirs, without their permission.<\/p>\n<p>\u201cNVIDIA has admitted training its NeMo Megatron models on a copy of The Pile dataset. Therefore, NVIDIA necessarily also trained its NeMo Megatron models on a copy of Books3, because Books3 is part of The Pile,\u201d the complaint reads. <\/p>\n<p>\u201cCertain books written by Plaintiffs are part of Books3 \u2014 including the Infringed Works \u2014 and thus NVIDIA necessarily trained its NeMo Megatron models on one or more copies of the Infringed Works, thereby directly infringing the copyrights of the Plaintiffs.\u201d<\/p>\n<h2>Direct Infringement Damages<\/h2>\n<p>Relying on the same logic, the authors accuse the company of direct copyright infringement, noting that NVIDIA copied their books to use them for AI training purposes. Through the lawsuit, the rightsholders demand compensation in the form of actual or statutory damages.<\/p>\n<p>The class action lawsuit includes three authors thus far, but more may be added to the case as it progresses. NVIDIA has yet to respond to the allegations but in light of similar cases, it will likely oppose the claims and\/or argue a fair-use defense. <\/p>\n<p>Last month, OpenAI managed to <a href=\"https:\/\/torrentfreak.com\/court-dismisses-authors-copyright-infringement-claims-against-openai-240213\/\">\u2018defeat\u2019<\/a> several copyright infringement claims from book authors in a somewhat related \u201cBooks3\u201d lawsuit. However, the California federal court didn\u2019t review the direct copyright infringement claims in this case, which have yet to be argued in detail at a later stage. <\/p>\n<p><em>\u2014<\/em><\/p>\n<p>A copy of the class action complaint against NVIDIA, filed by the authors in a California federal court, is <a href=\"https:\/\/torrentfreak.com\/images\/nvidia-lawsuit.pdf\">available here (pdf)<\/a><\/p>\n<p>From: <a href=\"https:\/\/torrentfreak.com\/\">TF<\/a>, for the latest news on copyright battles, piracy and more.<\/p>\n<p class=\"wpematico_credit\"><small>Powered by <a href=\"http:\/\/www.wpematico.com\" target=\"_blank\" rel=\"noopener\">WPeMatico<\/a><\/small><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Starting last year, various rightsholders have filed lawsuits against companies that develop AI models. The list of complainants includes record labels, book authors, visual artists, even the New York Times. These rightsholders all object to the presumed use of their work without proper compensation. \u201cBooks3\u201d Many of the lawsuits filed by book authors come with [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":75897,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[308],"tags":[],"class_list":["post-75896","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-torrent"],"_links":{"self":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/posts\/75896","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=75896"}],"version-history":[{"count":0,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/posts\/75896\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/media\/75897"}],"wp:attachment":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=75896"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=75896"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=75896"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}