Anthropic released Claude Sonnet 5.5, which beats Opus 5.5 at half the price
Anthropic released its flagship on 22 September. On the 28th it released the cheap one, and the cheap one scores higher on the benchmark the flagship led with, at half the input price.
Claude Sonnet 5.5 lists at $2 per million input tokens and $10 per million output, which is exactly what Sonnet 5 cost, with cache reads at $0.20. It scores 70.6% on Terminal-Bench 4.0; Opus 5.5, six days old and priced at $4 and $20, scores 66.4%. On GDPval-AA the two are level, 1844 against 1846. Sonnet 5.5 takes text and images, holds a 1M token context, returns up to 128K tokens, and generates output more than 30% faster than Sonnet 5. Anthropic's sharper claim is about the bill rather than the rate card: up to 30% less per task, because it reaches the same answer in fewer tokens.
Why this one is different
The mid tier is normally the flagship with something taken away, released later and scoring lower. Here it landed six days after the flagship and beat it on the headline benchmark, at half the price, which makes Opus 5.5 hard to justify for most of the work people were about to give it. A company that cuts the price of its own week-old flagship by releasing something cheaper and better is not managing a product line. It is racing.
The flagship is six days old and already the expensive option.
How we got here
- 22 Sep 2026Claude Opus 5.5 at $4 and $20, and GPT-6 Sol and Luna at half the old tier price, ninety minutes apart.
- 26 Sep 2026OpenAI's revenue run rate nears $70bn, up more than 70% since the quarter began.
- 28 Sep 2026Anthropic's IPO prospectus targets a valuation above $2tn.
- 28 Sep 2026Sonnet 5.5: same price as Sonnet 5, above Opus 5.5 on Terminal-Bench.
What it does and does not mean
Both numbers are Anthropic's, measured on Anthropic's harness, and a single benchmark where the cheaper model wins does not make it the better model: Opus 5.5 still leads on the long-horizon work the benchmarks here do not capture, and the GDPval scores are close enough to be noise. The claim worth testing yourself is the token one, because it is the only figure that shows up on an invoice. What the fortnight establishes is a pattern. Three price cuts in seven days from two companies, each arriving within hours or days of the other, is what a market does when the product has stopped being scarce.