Meta’s Muse Spark 1.3 has overtaken Gemini 3.8 Flash on the current Artificial Analysis Intelligence Index—but that does not make it the automatic new default. The independently measured xhigh route scores 61 versus Gemini 3.8 Flash at 59; a limited-preview max route reaches 62. Gemini remains about 66% faster in output and has a lower standard API price. DeepSeek V4 Flash is cheaper again, offers open weights, and answers much sooner.
Released on 2 September 2026, Muse Spark 1.3 is Meta’s clearest attempt yet to turn its AI comeback into a product builders can route real work through. The useful question is not which launch slide wins. It is which combination of intelligence, speed, data policy, and agent harness fits your workload.
Three takeaways
- Muse now leads Gemini on one serious composite index. Artificial Analysis scores Muse Spark 1.3 max at 62 and xhigh at 61, ahead of Gemini 3.8 Flash high at 59. That is a benchmark-specific change, not proof of universal superiority.
- “Cheap” has three meanings. Muse xhigh costs slightly less per completed Intelligence Index task, Gemini costs less per standard API token, and DeepSeek V4 Flash has the lowest normal API rate. Muse Contributor is the cheapest token route here only if you accept Meta using prompts and completions for training.
- Terminal-Bench 2.1 and 4.0 are not a before/after line. They use different task sets and, in Google’s table, different evaluation provenance. Reading 89.4% and 19.1% as a 78% capability collapse is mathematically neat and methodologically wrong.
Independent Artificial Analysis measurements. Higher is better for index and speed; lower is better for latency and cost.
Gemini high: 59
Gemini: 302
DeepSeek: 1.50s
Gemini: $0.58
What Meta actually launched
Official: Meta says Muse Spark 1.3 is rolling out in Muse Code and the Meta Model API. The model is aimed at longer agentic workflows: juggling multiple work streams, recovering context, correcting gaps, asking for help when stuck, and confirming consequential actions. In Meta’s internal coding comparisons, it used roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2. Those efficiency numbers are vendor-reported; EyesTech has not rerun them. Meta launch post · evaluation methodology
The availability footnote matters. Existing reasoning modes are live, but Meta says max reasoning is coming after additional safety testing. Artificial Analysis describes the 62-point max route as a limited partner preview. The broadly available xhigh measurement is 61—still ahead of Gemini’s 59, but a different product state from the launch table’s max column.
Meta’s max-mode table reports 75.4% on DeepSWE 1.1, 59.4% on SWEAtlas CodeBase QnA, 88.8% on Terminal-Bench 2.1, 66.9% on OSWorld 2.0, and 98.1% on the 512K–1M MRCR long-context slice. The wins are not universal: Opus 5 remains ahead on GDPVal-AA v2, JobBench, OSWorld, and AutomationBench, while GPT-5.6 Sol leads DeepSearchQA and Meta’s internal instruction-following index in the same chart.
The Artificial Analysis Intelligence Index: Muse 62, Gemini 59
Observed: Artificial Analysis Intelligence Index v4.1.1 combines nine evaluations. Agents receive 34% of the weight, coding 24%, scientific reasoning 24%, and general capability 18%. Terminal-Bench 2.1 alone contributes 16% of the total. The evaluator estimates the composite’s 95% confidence interval at less than ±1 point, while warning that individual evaluation intervals can be wider. Methodology
Bars are normalized to the highest displayed score, 62. They are not percentages.
* Muse max is limited preview and has no announced public price. DeepSeek’s current index row lacks three evaluations included for newer models, so treat that comparison as directional.
This is the defensible meaning of “Gemini was dethroned”: both models arrived on the same date, and Muse moved two points ahead at the generally available xhigh setting and three at max on this specific independent index. It is not a claim that Muse wins every workload.
The counter-evidence sits on the same comparison page. Gemini 3.8 Flash high produces about 302 output tokens per second versus Muse’s 182, and its measured time to first token is 13.30 seconds versus 27.51 seconds. Gemini’s list rate is also lower at $0.75 per million fresh input tokens and $3.75 per million output tokens during its introductory period, compared with Muse standard at $1.25 and $4.25. Direct comparison
Dots sit on separate scales; they are not connected because these metrics cannot be collapsed into one honest number.
Gemini 59
Muse xh 61
Muse max 62
Muse 182
Gemini 302
Gemini 13.30
Muse 27.51
Pricing: Muse is efficient per task, not cheapest per token
Meta’s pricing page splits Muse into two commercial choices. Standard pricing excludes prompts and completions from model training. Contributor pricing is dramatically lower because the customer gives Meta permission to use that data to train future models. That is a data-governance choice, not a coupon code.
| Route | Cached input / 1M | Fresh input / 1M | Output / 1M | Key condition |
|---|---|---|---|---|
| Muse Spark 1.3 standard | $0.15 | $1.25 | $4.25 | Prompts/completions not used to train Meta models |
| Muse Spark 1.3 Contributor | $0.002 | $0.10 | $0.20 | Training permission; lower rate limits |
| Gemini 3.8 Flash introductory | $0.075 | $0.75 | $3.75 | Introductory through 31 Dec 2026 |
| DeepSeek V4 Flash off-peak | $0.007 | $0.22 | $0.66 | Outside specified weekday UTC peak windows |
| DeepSeek V4 Flash peak | $0.014 | $0.44 | $1.32 | 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri |
Meta rate card · Google rate card · DeepSeek rate card
For one illustrative long agent run with 1 million fresh input tokens and 200,000 output tokens, the token-only totals are $2.10 on Muse standard, $1.50 on Gemini, $0.704 on DeepSeek peak, $0.352 on DeepSeek off-peak, and $0.14 on Muse Contributor. The last number is commercially dramatic but unsuitable for confidential repositories or private customer data unless the organization has explicitly approved the training permission.
USD list price only. Shorter bars cost less. Excludes tools, retries, taxes, provider markup and human review.
At a rounded ₹95/$ illustration: ₹199.50, ₹142.50, ₹66.88, ₹33.44 and ₹13.30 respectively, before GST and payment/FX fees.
Artificial Analysis still finds Muse xhigh slightly cheaper per completed index task: $0.55 versus Gemini high at $0.58. That is possible because completed-task cost depends on token consumption as well as the rate card. DeepSeek remains far lower at $0.11 per index task in the available comparison, although its current row lacks three evaluations used for the newer models. The right procurement formula is:
accepted-artifact cost = tokens + tools + retries + latency + human verificationTerminal-Bench 2.1 is useful. Terminal-Bench 4.0 is not “broken.”
The most confusing launch-day comparison is Gemini 3.8 Flash: Google reports 89.4% on Terminal-Bench 2.1 and 19.1% on Terminal-Bench 4.0. Claude Opus 5 moves from 89.1% to 51.8% in the same published table. The numbers look incompatible because the benchmark name hides a major change.
89 tasks · Google self-computed · Terminus 2 default agent
66 changed tasks · official public leaderboard · different run provenance
Why 2.1 is good: it covers 89 real terminal tasks, runs agents in isolated environments, and uses executable verifiers to inspect the final container state. A confident narrative gets zero if the end state is wrong. It is public, repeatable, widely used, and one of the highest-weight components in Artificial Analysis’s current index. Terminal-Bench paper
Why 2.1 is no longer enough: a public benchmark gets easier to optimize against over time. Scores compress, tasks saturate, harness builders learn the failure modes, and data leakage becomes harder to rule out. A strong 2.1 score is still evidence of terminal competence; it is weaker evidence of unseen-task generalization.
What 4.0 fixes: Terminal-Bench’s maintainers describe it as a continuous benchmark. Version 3.0 replaced the old task set. Version 4.0 then fixed 19 tasks, removed eight—including saturated tasks and tasks with public solutions—and set a flat eight-hour timeout to reduce infrastructure noise. The current set contains 66 tasks. That maintenance makes the benchmark harder to game and the environment more stable. Maintainer explanation · repository
Why 4.0 can “not work well” for a launch comparison: it measures the model inside an agent, not the naked API. With only 66 expensive, long-horizon tasks, agent implementation, tool policy, refusals, output-token limits, resource configuration, and retry behavior can dominate. Google’s evaluation document says the 2.1 Gemini score was self-computed in Terminus 2, while its 4.0 score was copied from the official leaderboard’s best listed thinking level. That is not a matched A/B run. Google methodology
Muse has no official 4.0 result in Meta’s launch table. Until the same Muse and Gemini API checkpoints run through the same agent, tools, time budget, task release, and number of trials, Terminal-Bench 4.0 should answer “how did this complete agent stack perform?”—not “which base model is smartest?”
What X adds—and what it cannot prove
Mark Zuckerberg framed the release as a major coding and agentic jump, pointed readers to Muse Code and the API, and said Muse open weights are next. That is useful roadmap evidence, but still an executive launch claim.
Mark Zuckerberg on X · 2 Sep 2026
Artificial Analysis’s launch thread is more valuable for the purchasing decision because it puts the 62/61 index result beside task-level cost and identifies max as a limited preview.
Artificial Analysis on X · 3 Sep 2026
Field posts provide a different kind of signal. In one public video comparison, @OmedVibeCodes criticized both Muse and Gemini and said Gemini required a second attempt to produce a working scene. The post is worth seeing because it challenges launch-day certainty. It is not a benchmark: there is no public repository, controlled prompt log, fixed time/token budget, or artifact-level rubric.
@OmedVibeCodes on X · 3 Sep 2026
The useful pattern is not “X says the model is good” or “X says it failed.” It is that agentic systems need inspectable artifacts. A video can reveal a broken interaction or awkward design. It cannot quantify reliability without the prompt, code, logs, tests, retries, and acceptance rule.
Who should care
| Workload | Best starting candidate | Reason |
|---|---|---|
| Fast multimodal chat or interactive UI | Gemini 3.8 Flash | Higher measured output speed, lower TTFT and broader audio support |
| Difficult agentic work where task success dominates | Muse Spark 1.3 xhigh | Narrow index lead and strong cost per completed index task |
| Low-cost API automation with non-peak scheduling | DeepSeek V4 Flash | Far lower list rate and 1M context |
| Private local or controlled deployment | DeepSeek V4 Flash | MIT-licensed open weights; substantial hardware requirement |
| Disposable public-data experiments | Muse Contributor | Very low rate if training permission is acceptable |
| Confidential code or customer data | Muse standard, Gemini paid, or governed DeepSeek route | Avoid Contributor until legal/security owners approve training use |
What this changes
Muse Spark 1.3 makes Meta credible in the fast agentic-model tier again. Its xhigh route clears Gemini 3.8 Flash on the current Artificial Analysis index and nearly matches Gemini’s cost per completed index task despite higher list prices. Meta has also made a strategically aggressive bet with Contributor pricing.
But the release does not produce one obvious winner. Gemini still owns the speed line and the lower standard rate. DeepSeek owns the normal price line and open-weight control. Muse owns a narrow independent intelligence lead and an unusually cheap training-eligible route. The correct result is routing, not coronation.
Verdict: test the stack, not the screenshot
Put Muse Spark 1.3 on your shortlist if your workload is long-horizon coding, browsing, document production, or multi-step tool use. Start with xhigh, because that is the available configuration with an independent 61-point result. Do not budget around max until Meta publishes access and pricing.
Run the same 20–50 representative tasks through Muse, Gemini, and DeepSeek with one agent harness. Record task pass rate, output and cached tokens, wall-clock time, retries, refusal rate, human corrections, and accepted-artifact cost. Keep consequential actions behind tests, diffs, screenshots, or human approval.
Limitations: EyesTech did not run the three APIs for this analysis. Meta’s launch results are vendor-reported; Artificial Analysis uses its own harness; the DeepSeek comparison lacks three current index evaluations; X examples are anecdotal; and prices can change. Recheck max-mode access, open weights, pricing, and a matched Terminal-Bench 4.0 run at 7, 30, and 90 days.
Updated 3 September 2026. Confidence: High on release, price, and cited measurements; Medium on cross-harness purchasing implications.
Get the EyesTech Signal: model launches move fast. We track the version, harness, price, and deployment constraint that actually changes the decision.
Related EyesTech analysis: Claude Fable 5.1 · DeepSeek V4 vs Kimi K2.6
