Pyyan / News / 4 September 2026

SafetyOpenAI · Gray Swan

One hidden attack in twelve still gets through OpenAI's safest model

8.5%of 1,810 hidden injection attacks succeeded

An independent red team got 8.5% of its hidden prompt injections past GPT-6 Astra. That is the good number. The model before it let through 27%.

Gray Swan ran 1,810 curated attacks against Astra in tool use and computer use scenarios, the settings the model is actually sold for. 8.5% succeeded, against 27.0% for GPT-5.6 Sol. On OpenAI's own GPT-Red evaluation, defence against indirect prompt injection rose from 96.23% to 99.79%, and direct attacks on the instruction hierarchy are saturated at 99.99%. The same page records that Astra is the first model OpenAI has rated Critical for cybersecurity capability, meaning it can find unknown vulnerabilities and build working exploits with little human help.

Why this one is different

Most safety numbers are published because they flatter. This set is published with the failure rate in it, and the failure rate is the figure that matters for what the model is for. A model sold for autonomous computer use is a model that reads documents somebody else wrote. A wrong answer one time in twelve is an annoyance. An attacker's instruction obeyed one time in twelve is a different category of thing.

8.5% is the improved number. The model before it was 27%.

How we got here

  1. 10 Aug 2026GPT-5.6-Cyber ships to a vetted tier after training raises exploit chain completion from 1.5% to 95%.
  2. 1 Sep 2026OpenAI rates Astra Critical for cybersecurity under its own Preparedness Framework, the first model it has so rated.
  3. 3 Sep 2026Astra ships. Daybreak members first, and disabled by default for enterprise customers.
  4. 3 Sep 2026Booz Allen publishes an index showing an ordinary attack harness closing a 67 point gap between models.
  5. 4 Sep 2026The numbers behind the Critical rating are published. The residual is 8.5%.

What it does and does not mean

None of this says Astra is unsafe to use. 8.5% against 27.0% is a real improvement, it was measured by somebody outside the company, and the deployment carries safeguards to match: checkpoint encryption, monitoring of complete reasoning trajectories, and external misalignment monitoring on every tool using call. The figure is also measured against attacks built to work rather than against ordinary documents, so what it converts to in production depends on how many hostile inputs an agent actually meets, which nobody has published. What it does say is that the residual is not near zero, and it is being disclosed by the vendor rather than found by a journalist. Booz Allen's finding the day before was that tooling matters more than the model. This is the same point from the other end: the model's own score is a floor, not a ceiling.

Related

← All the news, newest first