研究人员对arXiv上的12,750篇论文进行了分析,通过校准检测器评估机器生成内容的规模 [1]。该检测器在0.4%的误报率下进行了校准,以2021-2022年ChatGPT发布前的论文作为基准 [1]。
分析涵盖了2023年1月至2026年7月期间发布的论文 [1]。结果显示,最新完整季度中约32%的论文呈现机器写作特征,而2026年初的峰值接近39% [1]。这一比例自ChatGPT发布以来在多个领域显著上升 [1]。
不同学科的检测结果差异巨大 [1]。计算机科学领域AI写作占比最高,达到约65%;而数学领域最低,仅约0.7% [1]。研究指出,数学领域的低比例与该学科文本结构的差异导致检测困难有关 [1]。
研究团队同时指出了该检测方法的多项局限性 [1]。检测器无法区分经过轻度编辑的文档与完全由机器生成的文档,因此报告的数字应理解为机器化写作流行度的下限 [1]。此外,该方法在覆盖率、控制样本量等方面也存在制约 [1]。
Researchers have analyzed 12,750 papers from arXiv using a calibrated detection system to assess the prevalence of machine-generated writing in academic submissions [1]. The findings reveal that approximately one-third of the most recent papers exhibit characteristics of AI-written content [1].
The detector was calibrated at a 0.4% false positive rate, using papers published between 2021 and 2022—before ChatGPT's release—as a baseline [1]. The study examined submissions spanning from January 2023 through July 2026 [1]. In the latest complete quarter analyzed, the proportion of papers flagged for AI writing reached approximately 32%, with the peak approaching 39% in early 2026 [1].
The prevalence of machine-generated content varies significantly across disciplines [1]. Computer science shows the highest adoption at roughly 65%, while mathematics registers the lowest at approximately 0.7% [1]. Researchers attribute the detection difficulty in mathematics to differences in text structure between human and machine writing in that field [1].
The research acknowledges important limitations of the detection methodology [1]. The detector cannot distinguish between lightly edited documents and fully generated ones, meaning the reported results represent a lower bound on the actual prevalence of AI-assisted writing [1]. Additional constraints include coverage limitations and sample size controls [1].