AlphaGo核心团队成员Thore Graepel在MIT Technology Review撰文指出,当今大语言模型(LLM)并不具备真正的推理能力,而是在进行快速的模式识别和token预测1。作者以AlphaGo在2016年3月首尔与李世石的比赛为例说明,该系统在第二盘下出的第37手在人类直觉评估中概率仅为万分之一,但正因为具有搜索和推理机制才能选中这步创意之手,最终以4-1获胜1。相比之下,语言模型虽然在ChatGPT面世后通过引入"思维链"方法来展示推理过程,但这些思维链仍然基于相同的token预测机制,往往是事后编造而非真正的推理1。
作者认为当今LLM存在三个核心缺陷:缺乏明确持久的认知状态、知识与推理混杂于神经网络权重中、思维链内容事后编造1。Graepel主张构建新型AI系统,应该维持代表已知、怀疑、排除事项和开放问题的认知状态,将推理定义为改变认知状态的行动序列,这样才能在医学、工程等高风险领域产生可信结果1。作者近期已离开Google DeepMind,并力推采用类似AlphaGo架构的新方法来实现这一目标1。
Thore Graepel, a core member of the AlphaGo team, has argued that today's large language models lack genuine reasoning capabilities and instead rely on rapid pattern recognition and token prediction.1 Writing in MIT Technology Review, Graepel contrasts this with AlphaGo's creative move 37 during its March 2016 match against Lee Sedol in Seoul, which the system selected through search and reasoning mechanisms despite the move having only a one-in-ten-thousand probability according to human intuition models.1 AlphaGo ultimately defeated Lee Sedol 4-1 in that series.1
Graepel draws on Daniel Kahneman's framework distinguishing System 1 (fast intuition) from System 2 (deliberate reasoning) to illustrate the gap between current LLM capabilities and true reasoning.1 Following ChatGPT's rise, the field introduced "chain-of-thought" methods to address concerns about insufficient language fluency, yet these techniques still operate within the same token prediction process.1 Graepel identifies three core shortcomings in today's LLMs: the absence of explicit, persistent cognitive states; the entanglement of knowledge and reasoning within neural network weights; and chain-of-thought outputs that are often retrofitted after the fact.1
Graepel advocates for AI systems that maintain cognitive states representing what is known, suspected, excluded, and questioned, treating reasoning as a sequence of actions that modify these states.1 He emphasizes that such structured approaches would produce trustworthy results in high-risk domains such as medicine and engineering.1 Graepel recently departed Google DeepMind to pursue this alternative approach inspired by AlphaGo's architecture.1
评论
还没有评论,欢迎留下第一条。