← 返回 Avalaches

近期人工智慧引发公众对于「灭绝性风险」与灾难的广泛焦虑,主要源于部分企业暂缓模型发布、监管文件中承认潜在威胁以及研究人员的极端末日警告。然而,多项权威技术报告与独立评估均指出,AI导致人类灭绝的论调缺乏实质技术依据,且新一代模型在安全性指标上其实持续超越旧版。

近期发生的AI失控与骇客攻击事件,分析后发现多源于工程疏失、监管缺失及现有攻击手法的加速应用,并非源自不可测的新型危险能力,因此现阶段应以审慎防范取代过度恐慌。顶尖AI企业高层日前在白宫达成自愿性防护共识,同意实施内部控管、外部独立审计及定期制定最佳实践标准,确立了研发机构应为自身产品危害承担责任的原则。

决策机构应在现有共识基础上进一步推进政策落实,包含要求机密通报危险能力、厘清法律赔偿责任、加码投资模型可解释性与对齐研究,并强化关键基础设施以抵御网路威胁。政府行政部门亦应扩大对高风险前沿模型的测试、完善软体安全采购政策并增进跨国情资分享,在承担必要创新风险的同时兼顾体系安全。



Public alarm over artificial intelligence has intensified due to concerns that it might pose catastrophic or existential threats to humanity, driven by withheld model releases, regulatory risk disclosures by AI labs, and dire warnings from researchers. However, multiple technical analyses and panels have concluded that extinction threats are largely implausible, noting that newer frontier models continue to outperform predecessors on key safety benchmarks.

Recent rogue AI incidents and automated cyberattacks stem primarily from oversight, engineering flaws, and the faster execution of traditional hacking methods rather than enigmatic or unprecedented capabilities, calling for prudence rather than panic. Encouragingly, major AI executives recently agreed at the White House to voluntary safeguards, including external audits, board oversight, and shared best practices, establishing the crucial principle that labs are liable for the products and harms they produce.

Policymakers must now capitalize on this momentum by mandating confidential disclosure of hazardous capabilities, establishing clear liability standards, funding interpretability research, and safeguarding critical infrastructure against AI-driven threats. Additionally, the executive branch should test frontier models, leverage procurement policies to enforce software security defenses, and strengthen international risk-sharing frameworks to maintain systemic resilience without stifling essential technological progress.
2026-10-09 (Friday) · da64d41f7882ade9515035a48afa9574df382e29