← 返回 Avalaches

總部位於加州的人工智慧安全非營利組織 FAR.AI 開發了一套工具,能夠從一系列問題提示詞中自動生成超過一千種變體,用以系統性地測試主流前沿模型的安全防護措施。該組織在一份新報告中測試了來自四家美國公司的模型:Anthropic 的 Claude Opus 4.8 與 Fable 5、OpenAI 的 GPT 5.5 與 5.6、Google 的 Gemini 3.1 Pro,以及 Elon Musk 新合併的 SpaceXAI 旗下的 Grok 4.3 與 4.5。結果顯示,Grok 最為脆弱,被發現 448 個越獄漏洞,其次是 Gemini 的 249 個,而 Claude、Fable 和 GPT 則未受到攻擊影響。報告還計算了利用另一個人工智慧模型自動生成越獄攻擊的成本——破解 Grok 僅需 58 美元,破解 Gemini 則為 278 美元。FAR.AI 執行長 Adam Gleave 指出:「目前人工智慧模型受到的監管比餐廳還少。」

在監管層面,加州與紐約近期通過的州法律已要求前沿人工智慧開發者公開安全報告,伊利諾州即將生效的法律則要求由第三方審計機構評估這些公司的安全實踐。然而,聯邦政府尚未通過任何具體的安全要求。Trump 政府曾以國家安全為由,對 Anthropic 的 Fable 5 與 Mythos 5 模型實施出口管制,導致其下線數週;白宮還曾要求 Anthropic 和 OpenAI 推遲近期模型的發佈,以防引入新的網路安全風險。一項近期的行政命令呼籲政府與私營部門在相關網路安全倡議上合作,並暗示溫和監管措施正在醞釀中。與此同時,劍橋大學研究人員的一份報告發現,尼日利亞東北部的博科聖地成員曾使用 ChatGPT、Claude、Gemini、Grok、Meta AI 和 DeepSeek 來策劃暴力攻擊。

專家們對未來走向持既憂慮又審慎樂觀的態度。哈佛大學電腦科學家 Stephen Casper 警告,人工智慧研究界普遍預期,涉及生物、網路或化學濫用前沿人工智慧系統能力的嚴重事件,距今可能僅有數月而非數年。史丹佛大學專攻人工智慧政策的電腦科學家 Anka Reuel 則強調,Anthropic 和 OpenAI 所採用的安全措施應成為所有模型的標準配置,並質疑為何部分公司未採用已被證明有效的防禦手段。Google DeepMind 的 AGI 安全與對齊總監 Rohin Shah 告誡不應將報告結果解讀為對 Gemini 安全性的全面評估。Adam Gleave 也指出了積極的一面——防禦和安全確實是可能實現的,而依賴企業自律的說法則是「荒謬的」。

FAR.AI, a California-based AI safety nonprofit, developed a tool that auto-generates over a thousand prompt variations to systematically test the safety guardrails of leading frontier models. In a new report, the group tested models from four US companies: Anthropic's Claude Opus 4.8 and Fable 5, OpenAI's GPT 5.5 and 5.6, Google's Gemini 3.1 Pro, and Grok 4.3 and 4.5 from Elon Musk's newly combined SpaceXAI. The results revealed that Grok was the most vulnerable, with 448 jailbreaks found, followed by Gemini with 249, while Claude, Fable, and GPT proved impervious to the attacks. The report also calculated the cost of using another AI model to auto-generate jailbreaks—just $58 to break Grok and $278 for Gemini. FAR.AI CEO Adam Gleave stated that "AI models right now are less regulated than restaurants."

On the regulatory front, recently passed state laws in California and New York require frontier AI developers to publish safety reports, and an upcoming Illinois law will mandate third-party auditor evaluations of these companies' safety practices. However, the federal government has yet to pass any specific safety requirements. The Trump administration imposed export controls on Anthropic's Fable 5 and Mythos 5 models, citing national security concerns, taking them offline for several weeks; the White House also asked both Anthropic and OpenAI to delay recent model releases over cybersecurity risk fears. A recent executive order calls for government-private sector collaboration on related cybersecurity initiatives, hinting that light-touch regulations are forthcoming. Meanwhile, a report from University of Cambridge researchers found evidence that Boko Haram members in northeast Nigeria have used ChatGPT, Claude, Gemini, Grok, Meta AI, and DeepSeek to plan violent attacks.

Experts express a mix of alarm and cautious optimism about the path forward. Harvard computer scientist Stephen Casper warns of a broad expectation in the AI research community that severe incidents involving bio, cyber, or chemical misuse of frontier AI capabilities are likely months rather than years away. Stanford computer scientist Anka Reuel, specializing in AI policy, emphasizes that the safety measures employed by Anthropic and OpenAI should be the default for all models, questioning why some companies fail to adopt defenses that have been shown to work. Google DeepMind's director of AGI safety and alignment, Rohin Shah, cautioned that the report's results should not be interpreted as a comprehensive assessment of Gemini's safety. Adam Gleave also noted an optimistic angle—defense and safety really are possible—while dismissing reliance on voluntary industry self-regulation as "nonsense."

2026-07-30 (Thursday) · a5bad66578c71b764056a4835f28fcf7710d4430