尽管人工智能在数学证明领域取得显著成就,但其在科学创新中仍存在根本性局限。Google DeepMind 研究员 Tom Zahavy 在 ICML 2026 发表论文《LLMs can't jump》指出,大语言模型缺乏科学发现所需的关键能力——溯因推理,即从新现象提出新假设、新概念的能力 [1]。皮尔士将推理分为三类:归纳(从输入和输出推导规则)、演绎(由规则和输入推导输出)和溯因(由规则和输出推导新输入或新公理)[1]。当前的大模型虽然在归纳和演绎推理上表现出众,但无法实现溯因所需的创造性跳跃,即从感官体验到新公理体系的飞跃 [1]。
在数学领域,AI 模型确实已展现出强大的证明能力。OpenAI 的 GPT-5.6 Sol Ultra 仅用 3 页证明解决了圈双覆盖猜想,而 Anthropic 的 Fable 5 找到反例推翻了有 87 年历史的雅可比猜想 [1]。王虹与 Joshua Zahl 在 2025 年完成的三维挂谷猜想证明则长达 127 页 [1]。
然而,Zahavy 以爱因斯坦的等效原理和电梯思想实验为例说明大模型的局限。在爱因斯坦 1907 年至 1915 年构建广义相对论期间,物理学界并不存在需要被"压缩"的海量异常数据 [1]。他的突破源于新概念的创造性提出,而非数据驱动。与此同时,陶哲轩在 2026 年国际数学家大会指出数学界正面临"证明消化"问题——大量 AI 证明排队等待检查、验证 [1]。
Artificial intelligence systems have demonstrated remarkable prowess in mathematical proof-finding, yet they remain fundamentally constrained in generating the kind of conceptual breakthroughs that define scientific innovation. Researchers at Google DeepMind argue that large language models cannot perform the abductive reasoning—the imaginative leap from observations to new hypotheses and axioms—that is essential for genuine discovery.[1]
Recent AI achievements underscore both the capabilities and limitations of these systems. OpenAI's GPT-5.6 Sol Ultra proved the circle double cover conjecture in just three pages, while Anthropic's Fable 5 found a counterexample that disproved the Jacobi conjecture, which had stood for 87 years.[1] Wang Hong and Joshua Zahl completed a proof of the three-dimensional Kakeya conjecture spanning 127 pages in 2025.[1] Yet these accomplishments, however impressive, represent what large models do best: applying existing logical frameworks to solve predetermined problems.
Tom Zahavy, a researcher at Google DeepMind, articulated this distinction in a paper titled "LLMs can't jump," presented at ICML 2026.[1] Drawing on Peirce's tripartite classification of reasoning—induction (extracting rules from inputs and outputs), deduction (applying rules to generate outputs), and abduction (inferring new inputs or new axioms from rules and outputs)—Zahavy illustrates that while language models excel at induction and deduction, they cannot execute abductive reasoning.[1] Einstein's development of general relativity between 1907 and 1915 exemplifies the kind of conceptual innovation at stake: during this period, the physics community possessed no reservoir of anomalous data demanding compression and explanation, yet Einstein conceived entirely new principles through imaginative thought experiments like the equivalence principle and the elevator paradox.[1] Large models lack this capacity to generate new conceptual frameworks from first principles.
The challenge extends beyond theoretical concerns to practical consequences. At the 2026 International Congress of Mathematicians, Terence Tao highlighted an emerging "proof digestion problem" in mathematics: a growing backlog of AI-generated proofs awaiting verification and validation.[1]