由麻省理工学院和斯坦福大学等机构联合创建的AI Observatory项目发现,主流AI公司发布的使用数据与实际用户行为存在显著差异2。该项目汇总分析了包含85,633个对话轮次的数据,来自24,521个对话,涉及7个真实数据集2。研究表明,Anthropic、OpenAI等AI公司只发布他们选择公开的部分数据1。
研究揭露了数据报告中的具体偏差幅度。当应用Anthropic的过滤方法到AI Observatory数据集时,48%的对话会被过滤掉2。在非工作相关对话中,健康和人际关系话题占44.2%,而Anthropic报告中仅为31.2%2;成人或非法内容在真实数据中占7.9%,对比Anthropic的2.1%2;骚扰和仇恨内容分别为27.5%与5.66%2;色情内容为16.7%对比2.4%2。这些差异反映了AI公司报告中遗漏的敏感行为类别。
用户在不同AI模型间表现出差异化的使用习惯。研究发现,用户倾向于将Anthropic用于编程任务,Gemini用于社交和角色扮演,而ChatGPT则主要用于作业辅导1。OpenAI 2025年的报告显示,消费者中仅有30%的使用与工作相关2。
A collaborative research initiative has uncovered significant discrepancies between how AI companies report their products being used and what actually happens in practice. The AI Observatory, created by researchers from MIT, Stanford, and the Data Provenance Initiative, analyzed real conversation data from 5,000 users interacting with 52 different AI models to expose these gaps 2. The findings suggest that major AI firms like Anthropic and OpenAI publish only selective data about their services, leaving policymakers working with an incomplete picture of how these tools are actually deployed 12.
The research examined 85,633 conversation turns across 24,521 distinct conversations from seven real datasets spanning 2023 to 2025 2. When researchers applied Anthropic's content filtering methodology to the Observatory dataset, 48 percent of conversations would have been excluded 2. The filtered-out content reveals a stark contrast: in non-work-related conversations, health and relationship topics accounted for 44.2 percent compared to just 31.2 percent in Anthropic's publicly reported figures, while adult or illegal content represented 7.9 percent versus 2.1 percent 2. Similarly, harassment and hate speech appeared in 27.5 percent of conversations against 5.66 percent in company reports, and pornographic content in 16.7 percent versus 2.4 percent 2.
Users demonstrate distinct preferences across platforms based on their needs. Anthropic's Claude is predominantly used for programming tasks, Gemini attracts users seeking social interaction and roleplay, while ChatGPT serves primarily as a homework tutor 1. Anthropic's Economic AI Index drew from 1 million Claude conversations, while OpenAI's report analyzed 1.5 million ChatGPT conversations 2. An OpenAI 2025 report indicated that only 30 percent of consumer usage relates to work 2. Regarding the absence of comprehensive data, a researcher stated: "We don't have a way of answering that question because that information is proprietary" 2.
评论
还没有评论,欢迎留下第一条。