Anthropic released Claude Haiku 5.5 at a tenth of Haiku 4.5's price
Haiku 4.5 scored zero on Terminal-Bench 4.0. Not a low number: zero, on every task. Haiku 5.5 scores 39.2%, and costs a tenth as much to call.
Claude Haiku 5.5 lists at $0.10 per million input tokens and $0.50 per million output, for prompts up to 100,000 tokens. Past that it is $0.50 and $2.50. Haiku 4.5 charged $1.00 and $5.00. Anthropic puts the saving at about 75% on average, partly offset by a new tokenizer that spends slightly more tokens on the same task. The capability numbers moved further than the price. GDPval-AA goes from 735 to 1620 Elo. OSWorld 2.1 goes from 15.7% to 72.4%. Humanity's Last Exam, without tools, goes from 10.2% to 45.9%. It is also the first Haiku with an adjustable effort setting, which until now was a thing only the larger models had.
Why this one is different
A cheap tier usually buys you a worse model that is good enough for easy work. This one crossed the line from unusable to useful on whole categories of task. A zero on Terminal-Bench means the old model could not complete a single terminal task end to end; 39.2% means it finishes a third of them. The comparison that matters is not against its predecessor, though, but sideways: at $0.10 it is cheaper than GPT-6 Luna and beats it on every benchmark Anthropic published, and it does about half of what Sonnet 5.5 does for a twentieth of the price.
The old cheap model scored zero. This one finishes a third.
How we got here
- 22 Sep 2026Claude Opus 5.5 and GPT-6 Sol and Luna, ninety minutes apart, both cheaper than what they replaced.
- 28 Sep 2026Claude Sonnet 5.5 scores above Opus 5.5 at half the input price.
- 29 Sep 2026GPT-6.1 Sol matches Astra on DeepSWE at a fifth of the cost.
- 7 Oct 2026Claude Haiku 5.5 at a tenth of Haiku 4.5, with Terminal-Bench going from 0.0% to 39.2%.
What it does and does not mean
Every number here is Anthropic's. The tiered pricing is the detail to watch on a real bill: a prompt of 100,001 tokens costs five times as much per token as one of 99,999, for the whole request, and agent loops drift over that line without anyone deciding to. The tokenizer change also means the per-token saving is not the per-task saving. What the release confirms is where the fight has gone. Four price cuts in sixteen days, and the cheapest tier is now the one making the largest capability jumps, because that is where the volume is.