Pyyan / News / 1 September 2026

SafetyOpenAI

OpenAI says a model of its own crossed the line it wrote to catch dangerous ones

91.5%of cyber jailbreaks refused, from 59%

OpenAI wrote a framework to identify models too dangerous to release as they are. On 1 September one of its own models tripped it.

OpenAI warns of risks in new 'Astra' AI model · ReutersReuters' report on the same announcement. It covers the risk warning described above rather than the model's capabilities.
Astra, after safeguards91.5GPT-5.6 Sol590100 %
Share of cyber jailbreak attempts refused, OpenAI's own testing. The gap is what the Critical rating forced before release.

Astra is the first model OpenAI has rated Critical for cybersecurity under its Preparedness Framework. During testing it found and chained two zero-day vulnerabilities nobody had asked it to look for, and scored full marks on ExploitBench. The rating forces extra safeguards before release: Astra now refuses 91.5% of cyber jailbreak attempts against 59% for GPT-5.6 Sol, and the strongest capabilities go to vetted partners only.

Why this one is different

Safety frameworks are usually published and then never bind anything, because no model ever quite reaches the threshold. This one caught something, and the company said so before shipping rather than after. The specific detail worth holding on to is the unprompted part: the zero-days were not the task. They turned up while the model was doing something else, which is a different kind of capability from passing a test designed to measure it.

The zero-days were not the task. They turned up while it was doing something else.

How we got here

  1. 2023OpenAI publishes the Preparedness Framework, with capability tiers up to Critical and a commitment not to ship at that tier without safeguards.
  2. 10 Aug 2026GPT-5.6-Cyber, trained for offensive security work, released only through the vetted Daybreak Red tier.
  3. 1 Sep 2026Astra becomes the first model OpenAI rates Critical on any axis, and the framework does something for the first time.

What it does and does not mean

The grader and the graded are the same company. OpenAI wrote the framework, ran the evaluation, decided the rating, chose the safeguards and will decide who gets access, and no outside body checked any step. A 91.5% refusal rate also means roughly one attempt in twelve still gets through. What it does show is a lab publishing a threshold in advance and then honouring it against its own flagship, which is the first time the exercise has cost anybody anything.

OpenAISecurityWeekfrom the source itself

Related

← All the news, newest first