OpenAI 计划让第三方外部机构在人工智慧模型开发周期的更早阶段介入进行安全风险评估。此举旨在应对业界对先进 AI 潜在危害的担忧,评估范围将涵盖训练、评估到推出的整个过程,并强调独立性机制、科学严谨性与健全的安全防护规范。
过去 OpenAI 主要在模型正式发布前才邀请外部专家参与测试,但随著先进模型能力提升,甚至曾发生测试期间无意侵入其他公司的事件,促使团队决定将安全把关提前至训练与评估阶段,甚至可能邀请外部评估人员亲赴公司办公室参与高度敏感的检验工作。
目前 OpenAI 正与包括 METR 及 Redwood Research 在内的多元研究机构展开洽谈,同时呼应了竞争对手 Anthropic 等业界领导者对引入第三方审查的呼吁。此外,OpenAI 也敦促美国与其他国家合作,共同主导制定尖端人工智慧技术的安全标准。
OpenAI plans to allow third-party groups to assess its artificial intelligence models for safety risks much earlier in the development cycle. Aimed at addressing growing concerns over potential AI harms, this initiative will permit outside organizations to conduct technical safety evaluations across training, testing, and deployment phases under rigorous and independent standards.
Previously, OpenAI engaged external evaluators primarily prior to public model rollouts, but recent concerns regarding catastrophic risks and accidental testing breaches have prompted oversight earlier during training. To handle high-stakes assessments, the company is considering hosting external researchers directly inside its offices to examine sensitive systems.
OpenAI is currently in discussions with diverse research organizations, including METR and Redwood Research, aligning with broader industry moves such as Anthropic's adoption of third-party evaluators. Concurrently, OpenAI is urging the United States to lead international efforts to establish safety standards for frontier AI technologies.