SafetyOpenAI · METR · Redwood Research
OpenAI will let outside assessors test models during training, not just before launch
4priority areas opened to outside assessment
OpenAI published its principles for third-party assessment. The change is when, not whether. Outside groups may now test models during training and evaluation, rather than only on a finished model before launch. Four areas are open to them: safety cases, critical safeguards, the capability evaluations behind the Preparedness Framework, and investigations of misalignment incidents. The most sensitive work may happen inside OpenAI's offices. It is in talks with METR and Redwood Research. No partner is named and no access terms are set, which is where Anthropic's version got to four days earlier by signing one.