Grok 4.6 raised its cached input price, which is the number most people never read
Everybody compares models on input and output price. For anything agentic, the number that decides the bill is the third one, and Grok 4.6 put it up.
Grok 4.6 has a 500K token context at $2 in and $6 out per million tokens, and cached input raised to $0.50 per million.
Why this one is different
Cache pricing is where long context actually gets paid for. An agent that re-reads a large context on every step pays the cached rate far more often than the fresh one, so a change there moves a real bill much more than a headline rate does. Anthropic moved the same lever in the opposite direction three weeks later, cutting Fable 5.1 cache reads to a quarter, and took up to 45% off an agentic run by doing it.
Cache pricing is where long context actually gets paid for.
How we got here
- 2024Prompt caching arrives across the major APIs, and long context stops being priced as if every token were new.
- 12 Aug 2026Grok 4.6 ships with cached input raised to $0.50.
- 1 Sep 2026Anthropic cuts Fable 5.1 cache reads to a quarter of the previous rate, moving the same lever the other way.
What it does and does not mean
A cache rate on its own tells you very little, because hit rates depend entirely on how an application is written, and two teams on identical pricing can see bills that differ several fold. Nobody outside xAI has published measured costs on 4.6. The point worth keeping is narrower: the comparison everybody makes, input against output, is not the comparison that determines what agentic work costs.