OpenAI本周发布了数百份声称解决世界级数学难题的解答,但这些成果未能满足数学咨询团体AGMAI制定的相关标准1。AGMAI在9月底发布的指导原则首项建议为"停止在专有模型上测试高级数学问题",然而OpenAI明确表示正在使用专有模型进行评估1。剑桥大学和伦敦国王学院的研究论文指出,OpenAI模型在将自然语言证明翻译为Lean形式化代码时存在至少两处差异,暴露了AI自动形式化过程中的可信度问题1。
在OpenAI发布的719份手稿中,仅10份包含了模型的推理过程,而其中只有42%的证明进行了形式化处理1。数学家Terence Tao在社交媒体上指出,AI提示器在初始目标"解决"后对更广泛的领域不感兴趣,无法充分理解AI输出1。Melanie Wood表示:"在发布时没有人类理解,现在工作才开始"1,强调数学界对AI模型解答缺乏人类理解的关切,这与AGMAI要求确保人类理解的原则相悖1。
OpenAI released hundreds of claimed solutions to mathematics problems this week, but the submissions have failed to meet standards established by the Advisory Group on Mathematics and AI (AGMAI) 1. Research papers from Cambridge University and King's College London identified discrepancies in how OpenAI's models translated natural language proofs into Lean code, exposing credibility concerns in AI's automated formalization process 1.
The mathematical community has raised significant concerns about the nature of these solutions. Of the 719 manuscripts released by OpenAI, only 10 included the model's chain-of-thought reasoning 1, while just 42 percent of the proofs underwent formal verification 1. A central tension exists between OpenAI's approach and AGMAI's guidance: the advisory group's primary recommendation issued in late September was to "stop testing advanced mathematics problems on proprietary models," yet OpenAI explicitly stated it evaluated results using proprietary models 1.
Leading mathematicians have voiced their criticism. Terence Tao commented on social media that AI prompt engineers lose interest in broader domains after the initial goal of "solving" is met and cannot fully understand AI outputs 1. Melanie Wood stated that there was no human understanding at the time of publication, and the work is only now beginning 1. The consensus reflects a core principle emphasized by the mathematical establishment: these AI-generated solutions lack the human comprehension that AGMAI insists must accompany any credible mathematical contributions 1.
评论
还没有评论,欢迎留下第一条。