{"id":81418,"date":"2025-02-01T09:00:29","date_gmt":"2025-02-01T09:00:29","guid":{"rendered":"https:\/\/www.cryptocabaret.com\/?p=81418"},"modified":"2025-02-01T09:00:29","modified_gmt":"2025-02-01T09:00:29","slug":"pirate-libraries-are-forbidden-fruit-for-ai-companies-but-at-what-cost","status":"publish","type":"post","link":"https:\/\/www.cryptocabaret.com\/?p=81418","title":{"rendered":"Pirate Libraries Are Forbidden Fruit for AI Companies. But at What Cost?"},"content":{"rendered":"<p><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2025\/02\/apples-600x606.jpg\" alt=\"apple\" width=\"300\" height=\"303\" class=\"alignright size-large wp-image-263257\" srcset=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2025\/02\/apples-600x606.jpg 600w, https:\/\/torrentfreak.com\/images\/apples-300x303.jpg 300w, https:\/\/torrentfreak.com\/images\/apples-148x150.jpg 148w, https:\/\/torrentfreak.com\/images\/apples.jpg 882w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\">Earlier this week, various rightsholder groups submitted their recommendations for the 2025 Special 301 Report. <\/p>\n<p>This annual overview, compiled by the U.S. Trade Representative, highlights countries that fail to live up to U.S. copyright protection standards.<\/p>\n<p>Various groups stressed the importance of copyright protection when it comes to new AI technologies. They argued that foreign governments should be mindful of potential copyright infringements. <\/p>\n<p>The Chinese government is called out, for example, for considering the introduction of a text and data mining (TDM) exception for AI. Other countries, including Japan, have already written AI exceptions into law. This raises concerns. Not just for copyright holders, but also for American tech giants. <\/p>\n<h2>Tech Companies &amp; Pirate Libraries<\/h2>\n<p>In the United States, explicit copyright exceptions for AI learning are non-existent. On the contrary, there are several high-profile lawsuits in the U.S. where tech companies including Meta, OpenAI, and Google are accused of copyright infringement. <\/p>\n<p>Rightsholders accuse these companies of training their LLMs (large language models) on content obtained from unauthorized sources, including pirate libraries. These repositories turned out to be a goldmine, as they contained a vast amount of text, free for the taking. The problem, however, is that copyright holders never gave permission to use it.<\/p>\n<p>The lawsuits will ultimately determine whether the tech companies are liable for copyright infringement, linked to this and other unauthorized use, or whether \u2018fair use\u2019 is a valid defense. <\/p>\n<p>It will take years before those cases are decided and, meanwhile, pirate libraries such as Z-Library, LibGen, and Anna\u2019s Archive are off limits. In countries where the law is more lenient or opaque, this might be an entirely different story. That could create a copyright schism with potentially far-reaching consequences.<\/p>\n<h2>DeepSeek \u2661 Anna\u2019s Archive<\/h2>\n<p>This week, hundreds of new articles were published on the latest AI model released by the Chinese company DeepSeek. This model isn\u2019t just accurate, it\u2019s also much cheaper to run, while significantly decreasing AI development costs.<\/p>\n<p>According to pundits, Deepseek poses a threat to American AI dominance and leadership. While early responses are often overblown, it shows that AI development is a serious, high stakes business. <\/p>\n<p>While DeepSeek\u2019s innovation doesn\u2019t stem from shadow libraries, the company did use them as key input. Recent publications have been less transparent about their data sources, but an earlier paper clearly mentions a reliance on Anna\u2019s Archive. <\/p>\n<p>\u201cWe cleaned 860K English and 180K Chinese e-books from Anna\u2019s Archive,\u201d a <a href=\"https:\/\/arxiv.org\/abs\/2403.05525\">DeepSeek VL paper<\/a>, published last March, states.<\/p>\n<\/p>\n<p><center><em>DeepSeek\u2019s prompted love letter to Anna\u2019s Archive<\/em><\/center><br \/><center><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2025\/02\/deepseek-anna.jpg\" alt=\"deepseek anna\" width=\"600\" height=\"336\" class=\"alignnone size-full wp-image-263331\" srcset=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2025\/02\/deepseek-anna.jpg 1056w, https:\/\/torrentfreak.com\/images\/deepseek-anna-300x168.jpg 300w, https:\/\/torrentfreak.com\/images\/deepseek-anna-600x336.jpg 600w, https:\/\/torrentfreak.com\/images\/deepseek-anna-150x84.jpg 150w\" sizes=\"auto, (max-width: 600px) 100vw, 600px\"><\/center><\/p>\n<h2>AI Teams Work with Anna\u2019s Archive<\/h2>\n<p>DeepSeek isn\u2019t alone in this. According to Anna\u2019s Archive, many AI teams, including those connected to large U.S. and Chinese companies, have reached out to the site, looking for fast access to data. <\/p>\n<p>Anna\u2019s Archive <a href=\"https:\/\/annas-archive.se\/llm\">offers to work with AI companies<\/a> in return for a generous donation or a data trade. While U.S. companies typically back off due to copyright concerns, other teams gladly work with the shadow library.<\/p>\n<p>\u201cWe\u2019ve provided about 20-30 companies\/teams with our entire dataset. It\u2019s the same data as on our torrents page, but they get access to high-speed SFTP servers.\u201d <\/p>\n<p>\u201cUsually, this is in exchange for a large monetary donation or, on occasion, in exchange for good datasets they acquired,\u201d <em>\u2018Anna\u2019s Archivist\u2019<\/em> adds, noting that all data they obtain is shared publicly. <\/p>\n<p>The shadow library provided copies of several redacted emails where companies requested access. We couldn\u2019t independently verify their authenticity, but they are worth sharing nonetheless. \u2018<\/p>\n<blockquote>\n<p><sub>\u201cWe are a research group from REDACTED, currently focusing on large language models (LLM) and in the process of data investigation. We are very interested in the high-quality resources you offer and would like to know more about the specifics.\u201d \u2013 Chinese company<\/sub><\/p>\n<\/blockquote>\n<p><center><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2025\/02\/annamail.jpg\" alt=\"email\" width=\"600\" height=\"391\" class=\"alignnone size-full wp-image-263322\" srcset=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2025\/02\/annamail.jpg 1376w, https:\/\/torrentfreak.com\/images\/annamail-300x196.jpg 300w, https:\/\/torrentfreak.com\/images\/annamail-600x391.jpg 600w, https:\/\/torrentfreak.com\/images\/annamail-150x98.jpg 150w\" sizes=\"auto, (max-width: 600px) 100vw, 600px\"><\/center> <\/p>\n<blockquote>\n<p><sub>\u201cWe saw your Twitter post about the 7.5M scanned Chinese academic non-fiction book collection you are offering for LLM training if that company contributes to digitizing them via OCR. We at REDACTED have state of the art OCR technology we can leverage and would like to discuss this with you. We are happy to share sample results and open source all the results, but would likely ask to keep our code\/pipeline proprietary.\u201d \u2013 US company. <\/sub><\/p>\n<\/blockquote>\n<h2>The \u201cForbidden Fruit\u201d<\/h2>\n<p>Faced with multi-million dollar lawsuits, large U.S. companies are no longer eager to work with Anna\u2019s Archive. However, AI teams in other countries are less reluctant, and that creates tension. <\/p>\n<p>The allure of shadow libraries for AI development is akin to the biblical forbidden fruit. Just as Adam and Eve were tempted by the tree of knowledge, AI developers are drawn to the vast troves of \u2018free\u2019 data within these unauthorized collections. <\/p>\n<p>Shadow libraries, filled with pirated works, offer the potential to train powerful AI models. However, like the original forbidden fruit, these shadow libraries come with a cost, at least for some. <\/p>\n<p>In the U.S., copyright laws and pressure from copyright holders, make AI companies hesitant to bite into this fruit, fearing legal repercussions. Reluctance could therefore place American AI development at a \u201cknowledge disadvantage\u201d. \u00a0 <\/p>\n<h2>Innovation: The AI Copyright Conundrum<\/h2>\n<p>Meanwhile, in countries with more lenient copyright exceptions for AI training, companies are free to indulge. They can feast on the knowledge offered by shadow libraries, potentially accelerating their AI capabilities and gaining a competitive edge. <\/p>\n<p>This has the potential to create a \u201ccopyright schism,\u201d where AI development in some countries surges ahead, fueled by readily available data, while others are held back by legal constraints.<\/p>\n<p>Without offering a value judgement, or engaging in too much hyperbole, this situation raises complex questions about the balance between protecting intellectual property and fostering innovation. <\/p>\n<p>Is it fair for some countries to have a knowledge advantage due to differing copyright laws? Could this lead to a global AI divide, where certain nations dominate the field due to their access to \u201cforbidden\u201d data?<\/p>\n<p>We don\u2019t have the answers to any of these questions. As highlighted earlier, rightsholders believe that more strict AI regulation worldwide is the answer. If AI companies want access, they can negotiate deals and pay for it. <\/p>\n<p>However, the shadow library understandably has a quite different take. <\/p>\n<p>\u201cThis could be a geopolitical argument for the West relaxing copyright rules. If the West wants to stay ahead in AI, archiving and distributing books should be made fully legal,\u201d \u2018Anna\u2019s Archivist\u2019 informs us.  <\/p>\n<p>From: <a href=\"https:\/\/torrentfreak.com\/\">TF<\/a>, for the latest news on copyright battles, piracy and more.<\/p>\n<p class=\"wpematico_credit\"><small>Powered by <a href=\"http:\/\/www.wpematico.com\" target=\"_blank\" rel=\"noopener\">WPeMatico<\/a><\/small><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Earlier this week, various rightsholder groups submitted their recommendations for the 2025 Special 301 Report. This annual overview, compiled by the U.S. Trade Representative, highlights countries that fail to live up to U.S. copyright protection standards. Various groups stressed the importance of copyright protection when it comes to new AI technologies. They argued that foreign [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":81419,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[308],"tags":[],"class_list":["post-81418","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-torrent"],"_links":{"self":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/posts\/81418","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=81418"}],"version-history":[{"count":0,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/posts\/81418\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/media\/81419"}],"wp:attachment":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=81418"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=81418"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=81418"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}