人工智能工具大规模抓取网站内容用于模型训练,正在打破互联网长达30年的内容流通生态。1根据Cloudflare的数据,超过50%的网络流量现已来自AI爬虫。1为了维护自身收益,越来越多网站开始采取措施阻止AI爬虫访问。1Cloudflare表示,从9月15日起,其管理的全球前10,000个网站中超过30%将在包含广告的页面默认屏蔽AI爬虫。1
这一趋势形成了一个恶性循环。随着更多网站部署反爬虫措施,AI模型能获取的训练数据质量持续下降。1最近一项研究发现,某些AI搜索工具所使用的源中,约六分之一本身是由AI生成的网站,进一步加剧了信息污染的风险。1此外,在Google推出AI总结功能后,维基百科英文版本的流量出现明显下降。1
Artificial intelligence tools have fundamentally disrupted the internet's three-decade-old traffic model by harvesting website content for model training rather than directing visitors back to source sites.1 In response, website operators are increasingly blocking AI crawlers to protect their revenue streams, a defensive move that threatens to degrade the quality of data available for AI model development.1
The scale of this disruption is substantial. Cloudflare, which manages over 30 percent of the world's top 10,000 websites, estimates that more than half of all web traffic now originates from AI bots.1 Beginning September 15, Cloudflare is implementing a default block on AI crawlers for pages carrying advertisements across its network of managed sites.1 This intervention reflects growing tension between content creators seeking to maintain audience engagement and AI developers requiring large datasets for training.
The consequence of these blocking measures is already visible in the declining quality of information available to AI systems. Recent research has found that approximately one-sixth of the sources used by AI search tools are themselves AI-generated websites, indicating contamination of training data.1 The impact on established platforms is evident: the English-language Wikipedia experienced notable traffic declines following Google's introduction of AI-powered summaries, suggesting that users increasingly bypass original sources in favor of AI-generated abstracts.1
评论
还没有评论,欢迎留下第一条。