美国法庭公开的解密文件显示,微软应用科学总监Brent Hecht在2024年1月将AI训练中的网络爬虫行为描述为"令人震惊的前所未有的盗窃"和"人类历史上最大的劳动力盗窃"1。这些文件揭露了OpenAI和微软在纽约时报等新闻机构发起的版权诉讼中的具体违规细节。
根据解密文件,OpenAI的中期训练数据集包含超过91,692份来自纽约时报、每日新闻和调查报道中心的作品副本1。衍生自Common Crawl的数据集仅从nytimes.com就包含超过200万份文件1。微软员工还策划了绕过纽约时报付费墙的方案,而OpenAI研究人员则删除了训练数据中的版权声明1。
微软的Copilot产品对出版商造成了实质威胁。内部文件显示,该"答案引擎"导致纽约时报域名的点击率相比传统必应搜索下降最多93%1。OpenAI ChatGPT负责人Nick Turley在内部沟通中指出,出版商面临"存在威胁",这些产品"基本上是替代性的"1。Project Mango数据集包含至少160,903件来自新闻出版商的独特作品副本1。微软首席执行官萨蒂亚·纳德拉在2024年早期的证词中表示,"任何付费内容都应该被想要使用它进行训练的任何人许可"1。
Newly unsealed court filings have exposed internal communications from Microsoft and OpenAI detailing how the companies conducted large-scale data collection from news publishers without authorization.1 In documents made public through litigation, Microsoft's applied science director Brent Hecht characterized the practice as a "shocking, unprecedented theft" and "the largest theft of labor in human history" in January 2024.1
The disclosed materials reveal the extent of content harvesting from major news organizations. OpenAI's intermediate training dataset contained over 91,692 copies of works published by The New York Times, The Daily News, and the Center for Investigative Reporting.1 A derivative dataset from Common Crawl alone included more than 2 million files sourced from nytimes.com.1 Additionally, a dataset called Project Mango contained at least 160,903 unique copies of works from news publishers.1 Internal OpenAI communications show that staff members devised schemes to circumvent The New York Times' paywall, and researchers removed copyright notices from training data.1
The documents also demonstrate the commercial impact of these practices. Microsoft's Copilot "answer engine" caused traffic to The New York Times domain to decline by as much as 93 percent compared to traditional Bing search results.1 Nick Turley, who leads ChatGPT at OpenAI, noted in internal communications that publishers faced an "existential threat," describing these AI products as "basically substitutional."1 Microsoft CEO Satya Nadella testified in early 2024 that "any paid content should be licensed by whoever wants to use it for training."1
评论
还没有评论,欢迎留下第一条。