{"id":76781,"date":"2024-05-02T09:00:30","date_gmt":"2024-05-02T09:00:30","guid":{"rendered":"https:\/\/www.cryptocabaret.com\/?p=76781"},"modified":"2024-05-02T09:00:30","modified_gmt":"2024-05-02T09:00:30","slug":"newspapers-sue-openai-for-copyright-infringement-and-fake-news-hallucinations","status":"publish","type":"post","link":"https:\/\/www.cryptocabaret.com\/?p=76781","title":{"rendered":"Newspapers Sue OpenAI for Copyright Infringement and \u2018Fake News\u2019 Hallucinations"},"content":{"rendered":"<p><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2024\/05\/newsprint.jpg\" alt=\"newsprint\" width=\"300\" height=\"191\" class=\"alignright size-full wp-image-249217\" srcset=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2024\/05\/newsprint.jpg 832w, https:\/\/torrentfreak.com\/images\/newsprint-300x191.jpg 300w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\">Starting last year, various rightsholders have <a href=\"https:\/\/torrentfreak.com\/category\/ai\/\">filed lawsuits<\/a> against companies that develop AI models.<\/p>\n<p>The list of complainants includes record labels, <a href=\"https:\/\/torrentfreak.com\/court-dismisses-authors-copyright-infringement-claims-against-openai-240213\/\">book authors<\/a>, visual artists, a <a href=\"https:\/\/torrentfreak.com\/authors-sue-nvidia-for-training-ai-on-pirated-books-240311\/\">chip maker<\/a>, and <a href=\"https:\/\/torrentfreak.com\/the-new-york-times-needs-more-than-imagined-fears-to-block-ai-innovation-240329\/\">news publications<\/a>. These rightsholders all object to the presumed use of their work without proper compensation.<\/p>\n<p>Keeping pace with the constant stream of legal paperwork is a challenge, but a complaint filed at a New York federal court yesterday deserves to be highlighted. In this case, eight major news publications are suing OpenAI and Microsoft for copyright infringement.<\/p>\n<h2>U.S. Newspapers Sue OpenAI and Microsoft<\/h2>\n<p>The New York Daily News, Chicago Tribune, Orlando Sentinel, Sun-Sentinel, Mercury News, Denver Post, Pioneer Press, and Orange County Register, claim that the AI companies used their publications to train and develop ChatGPT models without obtaining permission. <\/p>\n<p>In addition, ChatGPT can recall large parts of their copyright-protected articles, which effectively bypasses their paywalls. This has a direct effect on the newspapers\u2019 revenues, they argue. <\/p>\n<p>\u201cDefendants are taking the Publishers\u2019 work with impunity and are using the Publishers\u2019 journalism to create GenAI products that undermine the Publishers\u2019 core businesses by retransmitting \u2018their content\u2019\u2014in some cases verbatim from the Publishers\u2019 paywalled websites\u2014to their readers.\u201d<\/p>\n<h2>Training On and Reproducing Copyrighted Articles<\/h2>\n<p>The complaint alleges that the newspapers\u2019 articles are prominent parts of the training material for OpenAI\u2019s models. GPT-3, for example, has 175 billion parameters and includes the \u2018WebText2\u2019 and \u2018Common Crawl\u2019 databases that both contain material owned by the plaintiffs.<\/p>\n<p>This alleged unauthorized use remains ongoing, the newspapers claim, and it will likely continue in the future.<\/p>\n<p>\u201cOn information and belief, Microsoft and OpenAI are currently or will imminently commence making additional copies of the Publishers\u2019 Works to train and\/or fine-tune the next generation GPT-5 LLM,\u201d the complaint adds. <\/p>\n<p>The plaintiffs show that ChatGPT can reproduce content from copyrighted news articles when prompted. In addition, third-party services in the OpenAI store are specifically marketed to bypass their paywalls, they say.<\/p>\n<p>These tools include a custom GPT called \u201cRemove Paywall\u201d and a tool such as \u201cNews Summarizer\u201d, which promises to \u201csave on subscription costs\u201d and \u201cskip paywalls just using the link text or URL.\u201d<\/p>\n<\/p>\n<p><center><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2024\/05\/remove-paywall.jpg\" alt=\"remove paywall\" width=\"550\" height=\"356\" class=\"alignnone size-full wp-image-250960\" srcset=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2024\/05\/remove-paywall.jpg 1202w, https:\/\/torrentfreak.com\/images\/remove-paywall-300x194.jpg 300w\" sizes=\"auto, (max-width: 550px) 100vw, 550px\"><\/center><\/p>\n<p>OpenAI and Microsoft have previously argued that the use of copyrighted works to train its models falls under fair use. In addition, they called out the lack of specific copyright infringements by third parties. <\/p>\n<p>This lawsuit is likely to trigger similar defenses, but copyright infringement allegations are just part of the newspapers\u2019 complaint. <\/p>\n<h2>\u2018Fake News Hallucinations\u2019<\/h2>\n<p>The newspapers are not only concerned by the unauthorized use of their works; they also allege that the AI tools cause commercial and competitive injury by spreading false claims. <\/p>\n<p>The plaintiffs cite various examples where ChatGPT allegedly links dubious news reporting to their newspapers. <\/p>\n<p>\u201cAs if plagiarizing the Publishers\u2019 work were not enough, Defendants\u2019 products are often subject to \u2018hallucinations\u2019 where those products malign the Publishers\u2019 credibility by falsely attributing inaccurate reporting to the Publishers\u2019 newspapers.<\/p>\n<p>\u201cBeyond just profiting from the theft of the Publishers\u2019 content, Defendants are actively tarnishing the newspapers\u2019 reputations and spreading dangerous disinformation.\u201d <\/p>\n<p>One example is the spurious claim that disinfectants can cure Covid. While many newspapers reported on these claims, they didn\u2019t endorse them. <\/p>\n<\/p>\n<p><center><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2024\/05\/fakenews2.jpg\" alt=\"fake news\" width=\"600\" height=\"328\" class=\"alignnone size-full wp-image-250965\" srcset=\"https:\/\/www.cryptocabaret.com\/wp-content\/uploads\/2024\/05\/fakenews2.jpg 1188w, https:\/\/torrentfreak.com\/images\/fakenews2-300x164.jpg 300w\" sizes=\"auto, (max-width: 600px) 100vw, 600px\"><\/center><\/p>\n<p>These hallucinations dilute and injure the reputation of the newspapers, the complaint alleges. This claim comes on top of the various copyright infringement accusations for which they request compensation. <\/p>\n<p>Ultimately, the newspapers are not against Artificial Intelligence, but they do want OpenAI and Microsoft to pay for the content they use and, ideally, ensure that their reputations are not harmed in the process.<\/p>\n<p>\u201cThis lawsuit is about how Microsoft and OpenAI are not entitled to use copyrighted newspaper content to build their new trillion-dollar enterprises, without paying for that content.<\/p>\n<p>\u201cAs this lawsuit will demonstrate, Defendants must both obtain the Publishers\u2019 consent to use their content and pay fair value for such use,\u201d the newspapers conclude.<\/p>\n<p><em>\u2014<\/em><\/p>\n<p>A copy of the complaint, filed by the newspapers at the U.S. District Court for the Southern District of New York, is <a href=\"https:\/\/torrentfreak.com\/images\/newspapers-openai.pdf\">available here (pdf)<\/a><\/p>\n<p>From: <a href=\"https:\/\/torrentfreak.com\/\">TF<\/a>, for the latest news on copyright battles, piracy and more.<\/p>\n<p class=\"wpematico_credit\"><small>Powered by <a href=\"http:\/\/www.wpematico.com\" target=\"_blank\" rel=\"noopener\">WPeMatico<\/a><\/small><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Starting last year, various rightsholders have filed lawsuits against companies that develop AI models. The list of complainants includes record labels, book authors, visual artists, a chip maker, and news publications. These rightsholders all object to the presumed use of their work without proper compensation. Keeping pace with the constant stream of legal paperwork is [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":76782,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[308],"tags":[],"class_list":["post-76781","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-torrent"],"_links":{"self":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/posts\/76781","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=76781"}],"version-history":[{"count":0,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/posts\/76781\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=\/wp\/v2\/media\/76782"}],"wp:attachment":[{"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=76781"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=76781"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.cryptocabaret.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=76781"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}