研究人员发现了一种方法,可以从 OpenAI、Anthropic 和 Google 等先进 AI 模型中提取隐藏的「思考」过程。这种攻击手法是将加密的推理追踪输入到同模型较小、对齐程度较低的版本中,借此揭露其内部推理,甚至可能导致密码和 API 金钥等个人资讯外泄(该漏洞现已修复)。
这项发现也让「推理蒸馏」的争议浮上台面。研究显示,某些中国开源模型(例如 Moonshot AI 的 Kimi K3)的推理输出与美国的 Claude Opus 等模型惊人地相似。虽然这无法作为直接的因果证据,但加深了美国对于竞争对手可能利用蒸馏技术低成本复制其先进技术的担忧。
尽管在地缘政治上引发关注,但专家对于蒸馏技术在竞争中的实际影响看法不一。有人认为,蒸馏只能在有限程度上提升现有模型,且中国企业似乎已具备从头开发顶尖模型的实力;Meta 执行长祖克柏也指出,蒸馏是开源生态系统运作的重要原则,过度限制反而可能使美国处于劣势。
Researchers have discovered a method to extract hidden "thinking" processes from advanced AI models developed by OpenAI, Anthropic, and Google. The attack involves feeding encrypted reasoning traces to a smaller, less aligned version of the same model to reveal its internal reasoning, which previously could have leaked personal information like passwords and API keys (a vulnerability that has since been patched).
This discovery highlights the ongoing controversy surrounding "reasoning distillation." The research notes that certain Chinese open-weight models, such as Moonshot AI's Kimi K3, produce reasoning outputs strikingly similar to those of US-based models like Claude Opus. While not conclusive proof of causation, this exacerbates US concerns that competitors might be using distillation to efficiently copy advanced capabilities.
Despite the geopolitical tensions, experts are divided on the actual competitive impact of distillation. Some argue that distillation only enhances existing models to a limited extent and that Chinese companies already possess the expertise to build cutting-edge models from scratch. Furthermore, Meta CEO Mark Zuckerberg has defended distillation as a crucial principle of the open-source ecosystem, warning that restricting it could put the US at a disadvantage.