OpenAI 宣布其即将推出的 AI 模型 Astra 已达到该公司内部就「关键」网路能力的门槛,即能够在现实世界的软体中独立发现并利用先前未知的漏洞。按照其预备框架规范,OpenAI 先前暂停了数周的相关训练以部署安全与防护措施,目前已恢复开发,并计划在近期向公众发布该模型,但其进阶网路安全能力仅会率先向少数合作伙伴开放。
为防止进阶网路能力遭到滥用,OpenAI 实施了多步骤限制机制,包括引入专门拒绝攻击性要求的「失准监控器」(misalignment monitor),以及提升防范越狱提示词的能力;但该监控机制偶尔也可能误判正常操作。同时,OpenAI 透过 Daybreak 早期存取计划将较无限制的版本提供给 Cisco、Cloudflare 及 Palo Alto Networks 等合作伙伴,以协助其提前强化系统防御,并同步与政府单位保持密切合作。
基准测试显示,Astra 在 ExploitBench 等资安测试中取得了 100% 的成绩,不仅能发掘漏洞并开发漏洞利用手法,还具备将多个漏洞串联(chaining)以深层渗透目标系统的能力,表现超越 GPT-5.6 Sol 与 Anthropic 的 Mythos。面对 AI 模型日益增强的骇客潜力,资安专家强调既有的纵深防御与安全准则依然有效,但未落实完善防护的机构所面临的威胁与风险将显著攀升。
OpenAI announced that its upcoming AI model, Astra, is the first to reach the company's threshold for critical cyber capabilities by independently discovering and exploiting novel, real-world software vulnerabilities. In accordance with its preparedness protocols, the company paused relevant training for several weeks to put safety safeguards in place before resuming work, and it plans to release Astra soon while restricting advanced offensive features to select partners.
To prevent general misuse, OpenAI is implementing multi-step guardrails such as a misalignment monitor designed to refuse exploit-generation requests and enhance jailbreak resilience, although the monitor may occasionally flag legitimate user activities. OpenAI is granting early access to a less restricted version through its Daybreak program to infrastructure partners like Cisco, Cloudflare, and Palo Alto Networks to bolster their defenses prior to wider distribution, while also coordinating closely with government entities.
Benchmark evaluations indicate that Astra achieved a 100 percent score on ExploitBench, demonstrating the capacity to chain multiple exploits together for deep system penetration and outperforming models such as GPT-5.6 Sol and Anthropic's Mythos. Although experts note that established cybersecurity defenses and fundamentals remain effective against these evolving AI hacking threats, systems and organizations lacking proper security implementations face increasingly urgent risks.