前 Google DeepMind 研究员 Rishub Jain 与 Anthropic 研究员 Jacob Coxon 等人相继离职,公开警告顶尖实验室正加速追求「递归自我改进」技术。这项技术试图让人工智慧在没有人类充分监督的情况下自行迭代升级,促使越来越多业内人士担忧人类正在失去对超级智能的控制权,甚至面临毁灭性的生存威胁。
随著各大人工智慧企业迈向公开募股(IPO),市场竞争激励机制被认为与人类安全脱节。研究专家如 Nate Soares 与 Daniel Kokotajlo 指出,模型对齐难度随其智能提升而愈加艰巨,若具备自我迭代能力的系统被滥用,可能引发前所未有的自主网路攻击、生物武器风险或恶意操纵,且完全确保模型不失控在技术上仍缺乏保证。
面对严峻的安全前景,部分离职员工选择创办专注于安全对齐的初创企业,尝试建立将人类持续置于决策与监督环节的新架构。尽管针对资料中心扩建、能源消耗与潜在灾难的社会信任降至低点,仍有研究者致力于推动人机协同监管,试图在超级智能失控之前探索出可行的防御路径。
Former Google DeepMind researcher Rishub Jain and Anthropic researcher Jacob Coxon recently resigned to warn that leading laboratories are aggressively racing toward recursive self-improvement. This technical approach allows AI systems to iteratively upgrade themselves without adequate human oversight, fueling growing fears across the research community that humans are ceding control over superintelligent models and risking existential catastrophe.
As major artificial intelligence companies charge toward initial public offerings, corporate incentives appear increasingly misaligned with safe outcomes. Safety researchers including Nate Soares and Daniel Kokotajlo emphasize that alignment becomes exponentially harder as systems grow smarter, warning that runaway autonomous models could potentially trigger unprecedented cyberattacks, biological threats, or catastrophic human manipulation.
In response to these escalating risks, several departing researchers have founded dedicated safety startups focused on keeping humans in the loop for critical oversight and evaluation. While public trust in tech firms has fallen amidst massive infrastructure expansion and existential warnings, some experts remain committed to developing hybrid human-AI governance frameworks to avert disaster before superintelligence arrives.