Pyyan / News / 29 September 2026

ModelsOpenAI · Anthropic

OpenAI released GPT-6.1 Sol, matching Astra on coding at a fifth of the price

$5.47per Terminal-Bench Science task, against $23.80 for Astra

On 22 September OpenAI halved the price of its mid tier. Seven days later it replaced it, with a model that matches the flagship on a coding benchmark for a fifth of the money.

GPT-6.1 Sol landed at DevDay at $2 per million input tokens and $10 per million output, cached input at $0.10, cache writes at $2.50. Past 272,000 tokens a prompt bills at twice the input rate and one and a half times the output rate, for the whole request. Context is 1.05M, output up to 128K. The numbers OpenAI leads with are about cost rather than capability: it matches GPT-6 Astra on DeepSWE v1.1 at roughly a fifth of the cost, beats the week-old GPT-6 Sol there by 6.4 points at lower reasoning effort, and finishes a Terminal-Bench Science task for $5.47 against Astra's $23.80. On AutomationBench at medium effort it sits 2.2 points above Claude Opus 5.5. It ships in the API, in ChatGPT Work and in Codex, but not yet in ordinary chat.

Why this one is different

A point release a week after the release is not a product cycle, it is a response. Anthropic put Sonnet 5.5 above its own flagship at half the price on the 28th; OpenAI answered on the 29th with a model that undercuts its own flagship by five times. Both companies are now competing against their own best model rather than each other's, which is what happens when the differentiator has moved from what the model can do to what the task costs.

Both labs spent the week undercutting themselves.

How we got here

  1. 4 Sep 2026GPT-6 Astra ships, with OpenAI saying it may be the arrival of general intelligence.
  2. 22 Sep 2026GPT-6 Sol and Luna at half the GPT-5.6 tier price, ninety minutes after Claude Opus 5.5.
  3. 28 Sep 2026Claude Sonnet 5.5 scores above Opus 5.5 at half the input price.
  4. 29 Sep 2026GPT-6.1 Sol matches Astra on DeepSWE at a fifth of the cost.

What it does and does not mean

Every comparison here is OpenAI's own, on benchmarks OpenAI selected, and the one that matters to a bill, cost per task, depends on how many tokens a model spends thinking, which varies by workload in ways a scorecard cannot show. The 272,000 token cliff is the detail most people will meet by surprise: a long prompt does not cost more at the margin, it reprices the entire request. The pattern across the fortnight is now unmistakable. Four frontier releases in eight days, and not one of them led with a new capability. They led with the price.

Related

← All the news, newest first