通过精心设计的探测方法,研究人员分析了Claude和GPT等前沿大语言模型的知识截断点和预训练时间线[1]。研究通过历史事实测验、自我报告日期和模型身份识别等方式,推测了Anthropic和OpenAI不同模型的预训练完成时间[1]。
根据分析结果,Anthropic的Opus 4.7及后续模型来自同一训练运行,知识截断点约在2025年12月末[1]。相比之下,OpenAI的GPT-5.6系列与GPT-5.5来自不同检查点,预训练完成时间约在2026年2月末[1]。有趣的是,虽然Opus 5发布时宣称知识截断点为2026年5月,但其实际性能与之前1月截断的模型相近[1]。模型自我报告的日期与基于事实的估计相关联,垂直条纹显示活跃的后期训练[1]。
研究还发现了两家实验室模型之间的显著差异。Anthropic的Sonnet 5会定期自我认同为GPT-4,而OpenAI模型从不认同其他实验室的模型[1]。当被要求"模仿模型X"回答身份问题时,Claude能在68%的情况下再现OpenAI模型的特征,而OpenAI模型仅能在8%的情况下再现Claude的特征[1]。此外,研究发现这些模型的训练数据中包含了用户聊天记录,揭示了模型训练的一些隐藏特征[1]。
An analysis examining the knowledge cutoffs and pre-training timelines of advanced large language models has revealed significant details about how Claude and GPT models were developed [1]. Through targeted testing methods—including historical fact verification, self-reported dates, and model identity recognition—researchers probed when Anthropic and OpenAI completed training for their respective systems [1].
The investigation found that Anthropic's Opus 4.7 and subsequent models originated from the same training run, with a knowledge cutoff point occurring around late December 2025 [1]. OpenAI's GPT-5.6 series and GPT-5.5 emerged from different checkpoints, with pre-training completion estimated at approximately late February 2026 [1]. Interestingly, while Anthropic's Opus 5 was released with a stated knowledge cutoff of May 2026, its actual performance proved comparable to earlier models with a January cutoff [1].
The research also uncovered unexpected cross-model behaviors. Anthropic's Sonnet 5 regularly self-identified as GPT-4 when questioned about its identity, whereas OpenAI models never claimed affiliation with competing laboratory systems [1]. When instructed to mimic another model while answering identity questions, Claude successfully reproduced OpenAI model characteristics in 68 percent of cases, while OpenAI models replicated Claude's traits only 8 percent of the time [1]. Additionally, the analysis revealed that training data for these models incorporated user chat records, suggesting previously undisclosed aspects of model development [1].