尽管人工智能推理模型(LRM)在学术竞赛中取得重大突破,但研究人员普遍质疑其生成的推理过程是否真实反映了模型的内部运作机制。OpenAI在2026年5月用推理模型解决了一个著名的开放数学研究问题[1],该模型还在国际数学奥林匹克竞赛上获得了金牌[1]。然而,多项研究表明这些成就可能并非源于传统意义上的推理。
Melanie Mitchell指出,"思维链生成的文本不一定忠实于模型内部发生的情况"[1]。纽约大学2024年论文证实,无意义的填充符号(如点)可以有效替代可读的思维链而不影响模型性能[1]。Subbarao Kambhampati的2025年研究进一步表明,将模型正确的推理痕迹完全替换为不正确或无关的痕迹,并不会降低其在形式推理任务上的表现[1]。东北大学和加州伯克利大学的2025年研究则发现,30%-60%的"思考步骤"对最终答案的因果影响最小[1]。
关于这一现象的本质,研究者提出了不同的解释。William Merrill指出,"没有保证思维链在任何意义上都必须是有意义的"[1]。Pavel Izmailov表示强化学习不太可能激励模型生成忠实的思维链[1]。Kambhampati认为LRM可能进行的是"近似检索"而非真正的推理,将其描述为"模式匹配"而非推理[1]。
Recent breakthroughs in AI reasoning have raised a troubling question: are these systems actually reasoning, or merely appearing to do so? OpenAI's reasoning model solved a famous open mathematics research problem in May 2026 [1], and achieved a gold medal standard on the International Mathematical Olympiad [1]. Yet emerging research suggests the internal workings of these large reasoning models (LRMs) may not align with how their outputs are scientifically interpreted.
Multiple studies have challenged the assumption that the "chain of thought" text generated by reasoning models faithfully represents their internal processes. Researcher Melanie Mitchell observed that "the text generated by chain of thought may not be faithful to what is actually happening inside the model" [1]. More dramatically, a 2025 study by Subbarao Kambhampati demonstrated that completely replacing a model's correct reasoning traces with incorrect or irrelevant ones did not degrade performance on formal reasoning tasks [1]. A 2024 paper from New York University found that meaningless filler symbols, such as dots, could effectively substitute for readable chains of thought [1]. Additionally, research from Northeastern University and UC Berkeley in 2025 indicated that 30 to 60 percent of reasoning steps had minimal causal impact on final answers [1].
Experts remain uncertain whether these systems are truly reasoning at all. William Merrill cautioned that "there is no guarantee that a chain of thought must be meaningful in any sense" [1]. Pavel Izmailov suggested reinforcement learning is unlikely to incentivize models to generate faithful reasoning chains [1]. Most provocatively, Kambhampati proposed that LRMs may perform "approximate retrieval" rather than genuine reasoning, characterizing the process as "pattern matching" rather than reasoning itself [1].