OpenAI在9月8日宣布解决纳维-斯托克斯问题后,剑桥大学研究团队发现其自然语言证明与形式化Lean代码版本存在重要差异1。具体来说,关键的Lemma 8.6部分在两个版本中的约束条件不匹配——自然语言版本要求某个值低于m+4,而Lean代码版本则要求低于m+51。
这一发现凸显了AI生成数学证明的可靠性问题1。研究团队用两周时间才识别出真实差异,这与OpenAI声称的88小时生成时间形成鲜明对比1。本周OpenAI发布了722篇数学论文,其中仅部分附带Lean证明且未经手工检查1。剑桥大学和帝国理工学院的研究人员表示,这类AI自动形式化无法起到同行评审的作用1。Kevin Buzzard认为:"我有信心纳维-斯托克斯问题已被正确解决,但对PDF中描述的证明远不那么有信心"1。
OpenAI announced on September 8 that it had solved the Navier-Stokes problem, but researchers at Cambridge University have uncovered a critical inconsistency between the natural language proof and its formal Lean code version 1. The discrepancy appears in Lemma 8.6, where the natural language version requires values to remain below m+4, while the Lean formalization requires them to stay below m+5 1.
The research team required two weeks to identify the actual difference, substantially longer than the 88 hours OpenAI claimed for generating the proof 1. Anders Hansen stated that "using such AI automatic formalization cannot serve the role of peer review" 1, while Kevin Buzzard remarked, "I am confident the Navier-Stokes problem has been correctly solved, but I am far less confident about the proof as described in the PDF" 1.
OpenAI released 722 mathematics papers this week, though only a portion included Lean proofs and none underwent manual review 1. This discovery underscores that automated formalization alone remains insufficient as a substitute for traditional peer review processes in mathematics.
评论
还没有评论,欢迎留下第一条。