美国加州北区联邦法院正式批准作家群体诉Anthropic的集体诉讼和解方案,和解金额达15亿美元 [1]。法官William Alsup在2025年6月23日的简易判决中明确划分了AI训练数据的两种获取方式:购买正版书籍仅用于模型训练被界定为"合理使用",而批量爬取盗版内容并永久存档则构成侵权行为 [1]。这一判决为全球AI企业敲响了警钟,标志着数据合规的分界线已正式确立。
法院对Anthropic设置了三条强制性要求:永久关停盗版素材数据库、保留完整的数据获取凭证以及接受第三方溯源审计 [1]。该案件带来的影响已波及全球范围,截至2026年5月,全球在审的AI版权相关诉讼共计112起,其中近八成集中在美国 [1]。国内AI版权案件在2026年暴涨216%,其中41%的纠纷与Anthropic案件性质相同 [1],反映出数据合规问题已成为业界共同的法律风险。
合规数据市场规模快速扩张。中国合规中文语料市场规模在2025年达到32.6亿元,同比增速超过40% [1];海外头部大模型合规数据授权成本已占总研发投入的35%,预计2026年末将上升至45% [1]。2026年5月,由中国大百科全书出版社等22家权威单位联合签署了《人工智能高质量语料库建设公约》[1],进一步推动了行业的规范化发展。
The U.S. District Court for the Northern District of California has approved a $1.5 billion settlement agreement in a class-action lawsuit brought by writers against Anthropic [1]. The ruling represents a significant decision on how AI companies may legally acquire training data, distinguishing between compliant and infringing practices.
In his summary judgment issued on June 23, 2025, Judge William Alsup established a clear legal framework for AI data acquisition [1]. Purchasing legitimate published books solely for training purposes constitutes "fair use," the court found, while bulk scraping of pirated materials combined with permanent storage of such content constitutes copyright infringement [1]. As part of the settlement, Anthropic must permanently shut down its database of pirated materials, maintain complete documentation of compliance measures, and undergo third-party audit of its data sourcing practices [1].
The implications of this decision extend globally, sounding an alarm for the AI industry worldwide. As of May 2026, there were 112 active AI copyright-related lawsuits under review globally, with nearly 80 percent concentrated in the United States [1]. Within China specifically, AI copyright cases surged 216 percent in 2026, with 41 percent of disputes sharing the same nature as the Anthropic case [1]. The settlement has also spurred the market for compliant Chinese-language training data, which reached 3.26 billion yuan in 2025, representing growth exceeding 40 percent [1]. For leading international large language model developers, licensing costs for compliant data now account for 35 percent of total research and development spending, with projections indicating this figure could rise to 45 percent by the end of 2026 [1]. In May 2026, twenty-two authoritative Chinese institutions, including the Chinese Encyclopedia Press, jointly signed the "Artificial Intelligence High-Quality Corpus Construction Covenant" [1].