人工智能模型在标准化智力测试中的表现参差不齐,在某些领域取得快速进展,但在基础认知能力上仍存在显著局限。2024年末,哥伦比亚大学团队研究发现最佳模型仅能解决18%的《纽约时报》Connections谜题1,到2025年初某些模型已能几乎完美地每次解决这类谜题1,展示了拼图游戏领域的进步速度。
然而,AI在多个关键认知领域仍表现不佳。空间推理被证实是模型最薄弱的环节,其3D心理旋转能力远低于人类水平1。谷歌与伊利诺伊大学厄本那-香槟分校2024年研究发现,模型容易在Knights and Knaves谜题的细微变化上失败1。Apple研究则表明大语言模型在Tower of Hanoi问题中,当盘子数或人数达到6及以上时开始出现错误1。此外,华盛顿大学、斯坦福大学和Allen Institute for AI的研究表明,大语言模型在逻辑网格谜题上同样面临困难1。这些发现揭示了机器与人类认知之间的根本差异,特别是在抽象推理和复杂逻辑推理等高阶思维能力上的鸿沟。
Artificial intelligence systems have demonstrated impressive gains on certain puzzle challenges, yet continue to struggle with cognitive tasks that humans navigate with relative ease. Researchers tracking AI performance across multiple intelligence tests have uncovered persistent weaknesses in spatial reasoning, abstract problem-solving, and logical inference—revealing fundamental gaps between machine and human cognition.1
Progress in some domains has been rapid. A Columbia University team found that in late 2024, the best-performing models could solve only 18 percent of The New York Times' Connections puzzles.1 By early 2025, however, certain models achieved near-perfect accuracy on these word-association challenges, demonstrating the speed at which AI can master pattern-matching tasks.1 Yet this trajectory does not hold across all cognitive domains. Spatial reasoning presents one of the starkest divides: AI performance on 3D mental rotation tasks lags substantially behind human capabilities.1
Failures in logic-heavy problems expose deeper limitations. Google researchers and the University of Illinois Urbana-Champaign documented how language models consistently fail when encountering minor variations of the Knights and Knaves puzzle, a task involving logical deduction.1 Apple researchers identified a performance ceiling in the Tower of Hanoi problem, with large language models beginning to fail once the number of disks reaches six or above.1 Similarly, work by researchers at the University of Washington, Stanford University, and the Allen Institute for AI revealed that language models struggle substantially with logic grid puzzles.1 These findings underscore how current AI systems, despite their fluency with language and pattern recognition, lack the robust reasoning mechanisms required for complex logical inference.
评论
还没有评论,欢迎留下第一条。