Pyyan / News / 3 September 2026

ModelsOpenAI

OpenAI said its new model may be the arrival of AGI, and shipped it switched off

100%ExploitBench, a perfect score

OpenAI's new model scores 100% on ExploitBench, a test of whether a system can find software vulnerabilities and chain them into a working attack. It arrives at enterprise customers switched off, and a named administrator has to decide to turn it on.

GPT-6 Astra was released on 3 September. OpenAI reports 97.6% on FrontierMath Tier 4, a set of research level mathematics problems, 99.9% on ARC-AGI-3, a reasoning test built to resist memorisation, and 100% on ExploitBench. It completes a computer use task in 40 minutes against 75 for the model before it. Pricing holds at $10 per million input tokens and $50 per million output, with a faster mode at double both. Access begins with companies in Daybreak, OpenAI's application only cybersecurity programme, and reaches paid ChatGPT plans, the API, Azure and Bedrock over the following days.

Why this one is different

Saturated benchmarks are not new. The sentence around them is. OpenAI's president said the model could eventually be seen as the arrival of artificial general intelligence, a claim the company has spent years declining to make about its own releases. The rollout carries information the announcement does not: a model that ships disabled, to a vetted programme first, is being released on different terms from every model before it.

A model that ships switched off, to a programme you have to apply to join.

How we got here

  1. Jul 2026An AI led cyberattack on Hugging Face. OpenAI delays its next model to add safeguards.
  2. 10 Aug 2026GPT-5.6-Cyber ships to a vetted tier after training raises exploit chain completion from 1.5% to 95%.
  3. 1 Sep 2026OpenAI rates Astra Critical for cybersecurity under its own Preparedness Framework, the first model it has so rated.
  4. 3 Sep 2026Astra ships. Daybreak members first, disabled by default for enterprise customers.
  5. 3 Sep 2026The same day, Sanders and Casar announce a bill to ban superintelligence outright.

What it does and does not mean

A saturated benchmark is the end of a measurement, not a measurement of general capability. FrontierMath Tier 4 at 97.6% and ARC-AGI-3 at 99.9% mean those two tests can no longer tell this model apart from the next one, which is a fact about the tests. Nothing in the three numbers describes what the model does on work that has no answer key. ExploitBench at 100% is the one with an obvious use, and it is the reason the rollout is shaped the way it is. What the day does show is that the company most careful never to say the word has now let its president say it, on the same day two members of Congress proposed making the thing he named a federal crime.

OpenAIAl Jazeerafrom the source itself

Related

← All the news, newest first