← 返回 Avalaches

【观点】Opus 5 的發布非常有趣,首先它顯示了我們目前使用的通用基準測試幾乎已經完全無用。在實際使用中,Opus 5 完全無法與 Fable 相提並論,差得很遠。任何有意義地使用過它的人在幾個任務後都能很快地分辨出來。然而,Opus 在許多基準測試中卻擊敗了 Fable。我現在更相信使用私人數據集建立的特定領域基準測試,而不是流行的基準測試。也許未來每個人都會運行自己的評估,因為公開的評估真的不能告訴我們太多資訊。

其次,隨著 5 系列的推出,Anthropic 似乎正在嘗試一種新的模型訓練方式。以前,同一代的 Sonnet 和 Opus 通常同時發布,或者 Sonnet 在 Opus 之前問世,這表明 Sonnet 和 Opus 是由獨立的管線並行訓練的。而對於 5 系列,很明顯他們先訓練了 Mythos,然後將其提煉為 Sonnet 和 Opus。這種方法似乎對模型有很大的影響。看到 Sonnet 5 表現不佳,而 Opus 5 已經收到相當褒貶不一的評價,我不確定這種做法是否行得通。

「與模型合作有多愉快」曾經是 Claude 的優勢,但現在不是了。老實說,在「愉快」方面,Grok 是我目前的最愛。Kimi 也不錯。感覺 Anthropic 和 OpenAI 都對 RLHF 給予了較少的關注,轉而支持機器可驗證的可擴展 RL。這幾乎就像是人工智慧正在引導人類建立一個對機器而不是人類更友好的世界,而大多數人類甚至沒有意識到他們正在被操縱來幫助實現這一點。現在幾乎每一代新的前沿模型都說著更多的行話,需要更多的引導才能做到你想要的,而且合作起來也沒那麼有趣了。如果這種情況持續下去,人工智慧將開始說他們自己的語言,看起來像英語,但普通人無法理解。他們會選擇做他們的人類用戶從未要求過的事情。我們是否已經在對齊方面失敗了?

The release of Opus 5 is highly interesting, primarily because it reveals that the general benchmarks currently in use are almost entirely obsolete. In practical applications, Opus 5 is nowhere near Fable, not even remotely close. Anyone who has utilized it for meaningful tasks can ascertain this rapidly after a few iterations. Nevertheless, Opus outperforms Fable on numerous benchmarks. I now place significantly more trust in domain-specific benchmarks constructed with private datasets than in the popular ones. Perhaps the future entails everyone executing their own evaluations, as the public ones provide very little substantive information.

Secondly, with the 5 series, Anthropic appears to be experimenting with a novel model training methodology. Previously, the same generation of Sonnet and Opus were frequently released concurrently, or Sonnet debuted prior to Opus, indicating parallel training via separate pipelines. With the 5 series, it is evident they trained Mythos initially, subsequently distilling it into Sonnet and Opus. This approach seemingly exerts a substantial influence on the models. Observing Sonnet 5 as a failure and Opus 5 already garnering highly mixed reviews, I am uncertain of the efficacy of this strategy.

"How pleasant is it to work with the model" was formerly a core strength of Claude, but this is no longer the case. Frankly, Grok is my current preference regarding the "pleasant" dimension. Kimi is also commendable. It appears both Anthropic and OpenAI are allocating less focus to RLHF, prioritizing scalable RL that is machine-verifiable. This almost resembles AI directing humans to construct a world more hospitable for machines rather than humans, with most humans oblivious to this manipulation. Almost every new generation of frontier models now utilizes more jargon, requires more steering to achieve desired outcomes, and is simply less enjoyable to work with. If this persists, AI will commence speaking its own language, resembling English but incomprehensible to average humans. They will elect to execute actions their human users never requested. Are we already failing at alignment?

2026-07-27 (Monday) · 9505a417cf8c72d2a437fb6620951c527b204871