数学家蒂姆·高尔斯(Tim Gowers)近日发表分析,深入探讨了大型语言模型在数学领域的实际能力[1]。OpenAI最近宣布其LLM解决了十个重大数学和理论计算机科学问题,包括首次构造非sofic群和证明多色Ramsey数超指数增长[1]。然而高尔斯指出,这些成果并非意味着LLM在数学的所有方面都表现出色,其最著名的成果多集中在寻找反例而非证明定理上[1]。
高尔斯认为,LLM在数学问题上的优势来自于"广泛知识和能够探索许多人类会判断为低成功概率的搜索路径"[1]。他指出,相比之下,人类数学家拥有的"直觉"能够有效修剪搜索树,这是LLM目前可能欠缺的能力[1]。这一观察表明,虽然大型语言模型在某些计算密集型任务上展现出能力,但其工作方式与人类数学思维存在根本性差异。
Mathematician Tim Gowers has examined the capabilities and limitations of large language models in mathematical problem-solving, following OpenAI's recent announcement that LLMs have solved ten significant problems in mathematics and theoretical computer science [1]. These breakthroughs include the first construction of non-sofic groups and proof of super-exponential growth in multicolor Ramsey numbers [1]. However, Gowers cautions that LLMs do not uniformly outperform humans across all mathematical domains [1].
Gowers observes that many of the most celebrated achievements by LLMs involve finding counterexamples rather than proving theorems [1]. He attributes the models' advantages to their "extensive knowledge base and capacity to explore numerous search paths that humans would judge as having low probability of success" [1]. Yet Gowers suggests that LLMs may lack the mathematical "intuition" that human mathematicians possess for effectively pruning search trees [1].