← 返回 Avalaches

這篇文章探討了人工智慧「推理」能力的真實性問題。大型推理模型(LRM)在數學奧林匹克競賽中獲得金牌、解決著名的開放數學問題,展現了令人矚目的成就。然而,多項研究同時揭示了這些模型的嚴重缺陷:它們在看似簡單的條件下會徹底崩潰,並且經常依賴表層捷徑而非真正的泛化推理來通過基準測試。這種矛盾的證據讓科學界和記者都感到困惑。

文章深入剖析了「思維鏈」(chain of thought)機制的本質。研究表明,LRM生成的推理痕跡既不忠實反映模型內部運作,也不一定對最終輸出有因果影響。紐約大學的研究發現,甚至用無意義的填充符號(如一串點)也能替代可讀的思維鏈而不影響表現。亞利桑那州立大學的Kambhampati團隊更證明,將正確的推理痕跡替換為錯誤或無關的內容,模型在形式推理任務上的表現並未下降。這些發現從根本上質疑了「思維鏈等同於推理過程」的假設。(關鍵數字:30, 60)

Kambhampati提出了一個替代假說:LRM本質上是在進行「近似檢索」,在龐大的訓練語料中進行模式匹配,而思維令牌的作用僅是填充上下文窗口,使模型更可能預測出類似推理的文本輸出。文章最終引用了1976年Drew McDermott提出的「一廂情願的助記符」概念,指出「推理」「思考」等術語可能正是這種認知陷阱的體現——我們因語言的擬人化力量而過度解讀了模型的行為。作者認為,在更清晰的科學解釋出現之前,應以類似「馬力」一詞的務實態度看待AI推理:承認其效用,但不必相信引擎蓋下真有馬蹄在奔跑。

This article investigates the contested nature of AI "reasoning" by examining the contradictory evidence surrounding large reasoning models (LRMs). On one hand, these models have achieved remarkable feats—solving open mathematical problems, winning International Mathematical Olympiad gold medals, and accelerating scientific research. On the other hand, rigorous studies from institutions like the Santa Fe Institute and Apple have shown that LRMs can suffer from complete accuracy collapse under simple perturbations and often rely on surface-level shortcuts rather than genuine generalized reasoning. The author, science journalist John Pavlus, frames his investigation around this persistent whiplash between triumph and failure.

At the heart of the debate lies the "chain of thought"—the stream of intermediate tokens that LRMs generate before producing a final answer. Multiple research teams have demonstrated that these reasoning traces are neither faithful representations of the model's internal processes nor causally necessary for correct outputs. A Northeastern University and UC Berkeley study found that 30–60% of thinking steps had minimal causal impact on answers. NYU researchers showed that meaningless filler tokens could substitute for readable chains of thought. Subbarao Kambhampati of Arizona State University characterizes these traces as "mumblings" and proposes that LRMs perform "approximate retrieval" across training data—a process closer to sophisticated pattern matching than stepwise logical reasoning—where thinking tokens serve mainly to load the context window favorably rather than narrate actual thought.

The article concludes by weighing the practical and scientific implications of this uncertainty. OpenAI's Sébastien Bubeck argues for a results-oriented perspective, emphasizing that models should be judged by what they accomplish rather than by mechanistic explanations. However, researchers like Melanie Mitchell and Tal Linzen counter that understanding whether a model is "right for the wrong reasons" matters deeply—particularly for building trust in non-verifiable domains and for discovering potentially superior training approaches. The author invokes Drew McDermott's 1976 concept of "wishful mnemonics" to frame the current discourse: terms like "reasoning" and "thinking" applied to AI may be begging the question, leading researchers and the public to anthropomorphize statistical processes. Pavlus ultimately adopts a pragmatic stance, likening AI reasoning to "horsepower"—a useful metaphor that should not be mistaken for literal horses under the hood.

2026-08-17 (Monday) · 7f0474774c6184ee57064ac12a4f6fac5bef5802