Meta's Muse Spark 1.3 does the same work with a quarter fewer tokens
The interesting number in Meta's release is not a benchmark. It is that the model does the same work with 25% fewer tokens and 20% fewer tool calls, which is the number that shows up on the bill.

Muse Spark 1.3 arrived on 2 September in Muse Code and the Meta Model API, which Meta calls its largest improvement yet on coding and agentic work. It scores 88.8 on Terminal-Bench 2.1 and 75.4 on DeepSWE v1.1, and 98.5 on the MRCR long context test between 256K and 512K. A maximum reasoning mode is held back pending further safety testing.
Why this one is different
Every lab claims a benchmark. Almost none of them publishes how much it costs to reach one. An agent that solves the task in twenty steps and one that solves it in fifteen score identically and bill very differently, and for anyone running agents at volume the second number decides whether the thing is affordable at all. Quoting the reduction alongside the score is the unusual part here.
Two agents can score identically and bill very differently.
How we got here
- 5 Aug 2026Muse Spark 1.2, Meta's agentic model with a terminal agent, at $1.25 in and $4.25 out per million tokens.
- 10 Aug 2026Muse Glimmer, 29.6B dense under Apache 2.0, the open counterpart to the proprietary Spark line.
- 2 Sep 2026Spark 1.3, quoting efficiency against 1.2 rather than only benchmark scores.
What it does and does not mean
These are Meta's own figures on Meta's own harness, and efficiency claims are unusually sensitive to how the test was set up: a different scaffold, a different tool set or a different retry policy moves the token count more than the model does. Nobody outside has reproduced them. What is worth taking from it regardless is that the industry has started competing on cost per completed task rather than score alone, which is the metric a buyer actually has.