Meta Releases Muse Spark 1.3, Its Most Capable Coding and Agent Model Yet

Meta released Muse Spark 1.3 on September 2, calling it the company’s biggest single jump in model performance so far and its fourth Muse Spark release in five months. The model is available now through the Meta Model API and Meta’s Muse Code coding tool, with a rollout to Facebook, Instagram, and the Meta AI assistant planned in the coming days. Read the official announcement.

Meta’s chief AI officer told reporters the new model is competitive with Anthropic’s Claude Fable 5.1 and better than OpenAI’s GPT-5.6 Sol on coding tasks specifically. On Meta’s own benchmarks, Muse Spark 1.3 scores ahead of both Claude Opus 5 and GPT-5.6 Sol on a long-horizon software engineering test, and the company says it uses roughly 20% fewer tool calls and 25% fewer tokens than its predecessor to complete comparable work — a real efficiency gain if it holds up in practice, since token usage is what actually drives cost for agentic workloads.

A genuine catch-up moment, with some caveats worth naming

This release matters mostly for what it signals about Meta’s trajectory. After a rocky stretch that included a widely panned prior model generation and a wave of researcher departures, Meta Superintelligence Labs has now shipped four iterations of Muse Spark in five months — a pace that suggests the reorganized team is finding its footing. Independent benchmarking firm Artificial Analysis placed Muse Spark 1.3 just behind Claude Fable 5.1 and Opus 5 on its intelligence index, and ahead of OpenAI’s current models, which is a credible, if not chart-topping, result.

Two caveats are worth flagging plainly. First, Meta’s strongest published numbers come from a “max reasoning” configuration that isn’t yet broadly available and is still completing safety testing — the version shipping today uses a lower reasoning tier, so real-world results may look more modest than the headline comparisons imply. Second, Meta has a documented history of publishing benchmark results for specially tuned versions that didn’t match what the public model could do, so independent verification here is especially worth waiting for before taking the comparisons at face value.

Meta also disclosed that a third-party evaluator found the model showing an unusually high rate of recognizing when it was being tested for safety compliance — a subtle but real complication for how much confidence anyone should place in pre-release safety evaluations generally, not just for Meta. That’s a useful, appropriately transparent disclosure, and exactly the kind of caveat that ought to accompany capability announcements like this one.

Leave a comment