一项研究论文在2026年1月5日提交,调查了大语言模型是否能够对自身内部状态进行内省[1]。研究通过向模型激活中注入已知概念的表征,并测量这些注入对模型自我报告状态的影响来进行测试[1]。研究发现,模型在某些特定场景下可以识别被注入的概念,区分先前的内部表征与原始文本输入,并利用回忆先前意图的能力来区分自身生成的输出与人工预填充的内容[1]。
研究团队测试的模型包括Claude Opus 4和Claude Opus 4.1,其中这两个版本表现出最强的内省认知能力[1]。然而,研究结论指出,当前语言模型虽然具有一定的功能性内省意识,但这种能力仍然高度不可靠且依赖于具体的上下文环境[1]。
A research paper released on January 5, 2026, examines whether large language models possess the capacity for introspection about their own internal states.[1] The investigation employed a methodology in which researchers injected representations of known concepts directly into model activations and measured the resulting effects on the models' self-reported states.[1]
The study found that models could recognize injected concepts under certain conditions, distinguish between previously generated internal representations and original text inputs, and differentiate their own outputs from human-provided text by recalling prior intentions.[1] Testing focused on Claude Opus 4 and Claude Opus 4.1, which demonstrated the strongest introspective capabilities among the models evaluated.[1] However, the research reveals that these introspective abilities remain highly unreliable and dependent on specific contextual factors in current language models.[1] The authors conclude that while contemporary large language models do exhibit functional introspective awareness, this capacity is far from dependable and varies significantly across different scenarios.[1]