这篇文章探讨了人工智慧「推理」能力的真实性问题。大型推理模型(LRM)在数学奥林匹克竞赛中获得金牌、解决著名的开放数学问题,展现了令人瞩目的成就。然而,多项研究同时揭示了这些模型的严重缺陷:它们在看似简单的条件下会彻底崩溃,并且经常依赖表层捷径而非真正的泛化推理来通过基准测试。这种矛盾的证据让科学界和记者都感到困惑。
文章深入剖析了「思维链」(chain of thought)机制的本质。研究表明,LRM生成的推理痕迹既不忠实反映模型内部运作,也不一定对最终输出有因果影响。纽约大学的研究发现,甚至用无意义的填充符号(如一串点)也能替代可读的思维链而不影响表现。亚利桑那州立大学的Kambhampati团队更证明,将正确的推理痕迹替换为错误或无关的内容,模型在形式推理任务上的表现并未下降。这些发现从根本上质疑了「思维链等同于推理过程」的假设。(关键数字:30, 60)
Kambhampati提出了一个替代假说:LRM本质上是在进行「近似检索」,在庞大的训练语料中进行模式匹配,而思维令牌的作用仅是填充上下文窗口,使模型更可能预测出类似推理的文本输出。文章最终引用了1976年Drew McDermott提出的「一厢情愿的助记符」概念,指出「推理」「思考」等术语可能正是这种认知陷阱的体现——我们因语言的拟人化力量而过度解读了模型的行为。作者认为,在更清晰的科学解释出现之前,应以类似「马力」一词的务实态度看待AI推理:承认其效用,但不必相信引擎盖下真有马蹄在奔跑。
This article investigates the contested nature of AI "reasoning" by examining the contradictory evidence surrounding large reasoning models (LRMs). On one hand, these models have achieved remarkable feats—solving open mathematical problems, winning International Mathematical Olympiad gold medals, and accelerating scientific research. On the other hand, rigorous studies from institutions like the Santa Fe Institute and Apple have shown that LRMs can suffer from complete accuracy collapse under simple perturbations and often rely on surface-level shortcuts rather than genuine generalized reasoning. The author, science journalist John Pavlus, frames his investigation around this persistent whiplash between triumph and failure.
At the heart of the debate lies the "chain of thought"—the stream of intermediate tokens that LRMs generate before producing a final answer. Multiple research teams have demonstrated that these reasoning traces are neither faithful representations of the model's internal processes nor causally necessary for correct outputs. A Northeastern University and UC Berkeley study found that 30–60% of thinking steps had minimal causal impact on answers. NYU researchers showed that meaningless filler tokens could substitute for readable chains of thought. Subbarao Kambhampati of Arizona State University characterizes these traces as "mumblings" and proposes that LRMs perform "approximate retrieval" across training data—a process closer to sophisticated pattern matching than stepwise logical reasoning—where thinking tokens serve mainly to load the context window favorably rather than narrate actual thought.
The article concludes by weighing the practical and scientific implications of this uncertainty. OpenAI's Sébastien Bubeck argues for a results-oriented perspective, emphasizing that models should be judged by what they accomplish rather than by mechanistic explanations. However, researchers like Melanie Mitchell and Tal Linzen counter that understanding whether a model is "right for the wrong reasons" matters deeply—particularly for building trust in non-verifiable domains and for discovering potentially superior training approaches. The author invokes Drew McDermott's 1976 concept of "wishful mnemonics" to frame the current discourse: terms like "reasoning" and "thinking" applied to AI may be begging the question, leading researchers and the public to anthropomorphize statistical processes. Pavlus ultimately adopts a pragmatic stance, likening AI reasoning to "horsepower"—a useful metaphor that should not be mistaken for literal horses under the hood.