全球领先的AI大模型正在依赖数十年来人类在开放知识平台上的无偿劳动积累。维基百科超级编辑Steven Pruitt编辑维基百科超过600万次,创建了3万多篇词条,但作为志愿者从未获得任何报酬[1]。约77%的维基百科词条编辑实际上来自仅占1%的活跃编辑者,这些被称为"超级编辑"的人承担了绝大部分工作[1]。ChatGPT、Gemini、Claude等AI系统都从维基百科这一"互联网的海马体"获取训练数据[1],但AI公司获取知识时缺乏透明度和相应的责任机制。
维基百科和其他开放知识平台的编辑贡献正逐渐被机器活动所取代。Amazon、Meta、Microsoft、Mistral AI和Perplexity等科技巨头已开始使用Wikimedia Enterprise的商业服务[1]。自2024年1月以来,维基百科下载多媒体内容所消耗的带宽增长了50%,至少65%的核心流量现已来自机器人[1]。与此同时,2025年部分月份维基百科的人类浏览量同比下降约8%[1]。为应对大模型对知识的侵蚀,2026年3月英语维基百科原则上禁止用大模型生成或改写词条正文[1]。
对于AI对维基百科内容的改编和使用,Pruitt对xAI的Grokipedia评价为"没有哪一点完全错,只是没有我写得那么准确"[1]。类似的知识众包模式也出现在其他领域,如ImageNet最初收录320万张图片、5247个语义类别,这些数据的分类工作依靠了Amazon Mechanical Turk平台上普通人的逐一标注[1]。这一现象反映出在AI时代,可追溯性和公共协作对于维护知识质量的重要性。
The training of advanced artificial intelligence systems relies heavily on decades of volunteer work accumulated on open knowledge platforms, with leading AI companies extracting this data while offering little transparency or compensation to the people who created it.[1]
Steven Pruitt, a Wikipedia supereditor, exemplifies this dynamic. Over his years of voluntary contribution, Pruitt has made more than 6 million edits to Wikipedia and created over 30,000 article entries without receiving any compensation.[1] Yet his work, along with that of roughly 1 percent of Wikipedia's most active contributors—who are responsible for approximately 77 percent of all article edits—now underpins the training datasets of major AI systems including ChatGPT, Gemini, and Claude.[1] When Pruitt reviewed xAI's Grokipedia, he observed that while the AI system was not fundamentally inaccurate, it lacked the precision of his original writing.[1]
The scale of AI companies' dependence on such publicly contributed knowledge has become increasingly visible. Amazon, Meta, Microsoft, Mistral AI, and Perplexity are among the firms using Wikimedia Enterprise's commercial services.[1] Meanwhile, the shift in traffic patterns reveals the growing presence of machine-generated requests: since January 2024, bandwidth consumed by multimedia downloads on Wikipedia has increased by 50 percent, with at least 65 percent of core traffic now originating from automated systems.[1] Simultaneously, human browsing of Wikipedia has declined by approximately 8 percent year-over-year during certain months in 2025.[1]
The challenge extends beyond Wikipedia to other platforms built on crowdsourced effort. ImageNet, which initially contained 3.2 million images classified into 5,247 semantic categories, relied on ordinary people using Amazon Mechanical Turk to manually categorize images one by one.[1] Recognizing these tensions, English Wikipedia has adopted a policy effective March 2026 that prohibits large language models from generating or rewriting article content.[1]