研究人員發現了一種方法,可以從 OpenAI、Anthropic 和 Google 等先進 AI 模型中提取隱藏的「思考」過程。這種攻擊手法是將加密的推理追蹤輸入到同模型較小、對齊程度較低的版本中,藉此揭露其內部推理,甚至可能導致密碼和 API 金鑰等個人資訊外洩(該漏洞現已修復)。
這項發現也讓「推理蒸餾」的爭議浮上檯面。研究顯示,某些中國開源模型(例如 Moonshot AI 的 Kimi K3)的推理輸出與美國的 Claude Opus 等模型驚人地相似。雖然這無法作為直接的因果證據,但加深了美國對於競爭對手可能利用蒸餾技術低成本複製其先進技術的擔憂。
儘管在地緣政治上引發關注,但專家對於蒸餾技術在競爭中的實際影響看法不一。有人認為,蒸餾只能在有限程度上提升現有模型,且中國企業似乎已具備從頭開發頂尖模型的實力;Meta 執行長祖克柏也指出,蒸餾是開源生態系統運作的重要原則,過度限制反而可能使美國處於劣勢。
Researchers have discovered a method to extract hidden "thinking" processes from advanced AI models developed by OpenAI, Anthropic, and Google. The attack involves feeding encrypted reasoning traces to a smaller, less aligned version of the same model to reveal its internal reasoning, which previously could have leaked personal information like passwords and API keys (a vulnerability that has since been patched).
This discovery highlights the ongoing controversy surrounding "reasoning distillation." The research notes that certain Chinese open-weight models, such as Moonshot AI's Kimi K3, produce reasoning outputs strikingly similar to those of US-based models like Claude Opus. While not conclusive proof of causation, this exacerbates US concerns that competitors might be using distillation to efficiently copy advanced capabilities.
Despite the geopolitical tensions, experts are divided on the actual competitive impact of distillation. Some argue that distillation only enhances existing models to a limited extent and that Chinese companies already possess the expertise to build cutting-edge models from scratch. Furthermore, Meta CEO Mark Zuckerberg has defended distillation as a crucial principle of the open-source ecosystem, warning that restricting it could put the US at a disadvantage.