《麻省理工科技评论》(MIT Technology Review)发布文章,通过一系列谜题测试展示了当前人工智能模型在空间推理、记忆适应性、抽象视觉推理、直觉和复杂逻辑等方面的优势与短板1。文章指出,尽管AI在部分基准测试上进步迅速,但在视觉谜题、细微变化识别和高复杂度逻辑任务上仍明显落后于人类1。作为最著名的基于谜题的基准测试之一,ARC-AGI要求模型从示例中推断抽象通用规则,进一步考验了AI的认知能力1。
2024年底,哥伦比亚大学科学家团队显示最佳AI模型仅能解决18%的《纽约时报》Connections谜题,而到2025年初,部分模型已能近乎完美地解决该谜题1。2024年,Google和伊利诺伊大学厄巴纳-香槟分校研究人员对Knights and Knaves谜题变体进行了模型训练和测试1。Apple研究人员发现,大语言模型能解决简单版本的汉诺塔问题和过河谜题,但当盘子或人数达到6及以上时开始出错1。此外,华盛顿大学、斯坦福大学和Allen Institute for AI研究人员观察到,大语言模型在逻辑网格谜题上同样表现挣扎1。该文章邀请读者亲自尝试这些曾难倒AI模型的题目,以比较人机认知差异1。
A recent publication by MIT Technology Review examines a series of puzzles designed to test the cognitive capabilities of current artificial intelligence models, revealing their strengths and shortcomings in spatial reasoning, memory adaptability, abstract visual reasoning, intuition, and complex logic 1. Despite rapid advancements in certain benchmark evaluations, AI continues to fall notably behind human performance in visual puzzles, the identification of subtle changes, and highly complex logical tasks 1.
In late 2024, a team of scientists at Columbia University showed that the best AI models could solve only 18% of The New York Times Connections puzzles, although by early 2025, some models had been able to solve them nearly perfectly 1. Researchers at Apple discovered that large language models (LLMs) can solve simple versions of the Tower of Hanoi and river crossing puzzles, but they start making mistakes when the number of disks or people reaches six or more 1. Researchers from the University of Washington, Stanford University, and the Allen Institute for AI observed that LLMs also struggle with logic grid puzzles 1. In 2024, researchers from Google and the University of Illinois Urbana-Champaign trained and tested models on variants of the Knights and Knaves puzzles 1.
ARC-AGI is one of the most famous puzzle-based benchmarks, requiring models to deduce abstract universal rules from examples 1. To compare the cognitive differences between humans and machines, the article invites readers to try these puzzles that previously stumped AI models 1.
评论
还没有评论,欢迎留下第一条。