Meta published Muse Spark 1.3 on 2 September 2026, a multimodal reasoning model aimed at long-running agentic, multi-agent and coding workflows. Its headline result is the highest DeepSWE v1.1 score in our tracked catalog.
The numbers
The output ceiling is the unusual specification. At 943,718 tokens it is built for generation runs that produce a great deal of material in one request — long agent traces, large refactors, extended documents — rather than for long prompts, which the 1M context window handles separately.
It accepts text, image, audio, video and file input, and Meta reports it completing coding work with roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2.
The caveat on that DeepSWE score
75.4% is the best DeepSWE v1.1 result we have recorded, ahead of GPT-6 Astra at 74.1% and DeepSeek V4.1 Flash at 74.2%. It is also, so far, the only coding benchmark Meta has published for this model on a version our recipe names.
That matters because DeepSWE is the generous end of the coding benchmark range. Across our catalog, DeepSWE scores cluster between 57% and 75%, while Terminal-Bench 4.0 scores run from 19% to 58%. A model measured on DeepSWE alone will rank above one measured on both, purely because the harder benchmark drags the second model's average down. We therefore do not publish a coding rating from a single benchmark family — Muse Spark 1.3's coding cell shows a labelled estimate until a second result lands.
This is not scepticism about the 75.4%. It is a refusal to let one strong result stand in for a whole category.
Where it fits
At $4.25 per million output tokens it sits mid-range: seven times the price of DeepSeek V4.1 Flash, a twelfth the price of GPT-6 Astra. If your workload is long agentic runs where the output ceiling and tool-call efficiency dominate, it is the specification to read closely.




