Meta’s Muse Spark 1.3 has overtaken Gemini 3.8 Flash on the current Artificial Analysis Intelligence Index—but that does not make it the automatic new default. The independently measured xhigh route scores 61 versus Gemini 3.8 Flash at 59; a limited-preview max route reaches 62. Gemini remains about 66% faster in output and has a lower standard API price. DeepSeek V4 Flash is cheaper again, offers open weights, and answers much sooner.

Released on 2 September 2026, Muse Spark 1.3 is Meta’s clearest attempt yet to turn its AI comeback into a product builders can route real work through. The useful question is not which launch slide wins. It is which combination of intelligence, speed, data policy, and agent harness fits your workload.

Three takeaways

  • Muse now leads Gemini on one serious composite index. Artificial Analysis scores Muse Spark 1.3 max at 62 and xhigh at 61, ahead of Gemini 3.8 Flash high at 59. That is a benchmark-specific change, not proof of universal superiority.
  • “Cheap” has three meanings. Muse xhigh costs slightly less per completed Intelligence Index task, Gemini costs less per standard API token, and DeepSeek V4 Flash has the lowest normal API rate. Muse Contributor is the cheapest token route here only if you accept Meta using prompts and completions for training.
  • Terminal-Bench 2.1 and 4.0 are not a before/after line. They use different task sets and, in Google’s table, different evaluation provenance. Reading 89.4% and 19.1% as a 78% capability collapse is mathematically neat and methodologically wrong.
EyesTech Model Reality Index · 3 Sep 2026
The launch-day result is a trade-off, not a crown

Independent Artificial Analysis measurements. Higher is better for index and speed; lower is better for latency and cost.

61
Muse xhigh Intelligence Index
Gemini high: 59
182
Muse output tokens/second
Gemini: 302
27.51s
Muse time to first token
DeepSeek: 1.50s
$0.55
Muse cost per index task
Gemini: $0.58

What Meta actually launched

Official: Meta says Muse Spark 1.3 is rolling out in Muse Code and the Meta Model API. The model is aimed at longer agentic workflows: juggling multiple work streams, recovering context, correcting gaps, asking for help when stuck, and confirming consequential actions. In Meta’s internal coding comparisons, it used roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2. Those efficiency numbers are vendor-reported; EyesTech has not rerun them. Meta launch post · evaluation methodology

The availability footnote matters. Existing reasoning modes are live, but Meta says max reasoning is coming after additional safety testing. Artificial Analysis describes the 62-point max route as a limited partner preview. The broadly available xhigh measurement is 61—still ahead of Gemini’s 59, but a different product state from the launch table’s max column.

Meta’s max-mode table reports 75.4% on DeepSWE 1.1, 59.4% on SWEAtlas CodeBase QnA, 88.8% on Terminal-Bench 2.1, 66.9% on OSWorld 2.0, and 98.1% on the 512K–1M MRCR long-context slice. The wins are not universal: Opus 5 remains ahead on GDPVal-AA v2, JobBench, OSWorld, and AutomationBench, while GPT-5.6 Sol leads DeepSearchQA and Meta’s internal instruction-following index in the same chart.

The Artificial Analysis Intelligence Index: Muse 62, Gemini 59

Observed: Artificial Analysis Intelligence Index v4.1.1 combines nine evaluations. Agents receive 34% of the weight, coding 24%, scientific reasoning 24%, and general capability 18%. Terminal-Bench 2.1 alone contributes 16% of the total. The evaluator estimates the composite’s 95% confidence interval at less than ±1 point, while warning that individual evaluation intervals can be wider. Methodology

Artificial Analysis · Index v4.1.1
Muse takes the narrow composite lead

Bars are normalized to the highest displayed score, 62. They are not percentages.

Muse Spark 1.3 max*62
Muse Spark 1.3 xhigh61
Gemini 3.8 Flash high59
DeepSeek V4 Flash max52

* Muse max is limited preview and has no announced public price. DeepSeek’s current index row lacks three evaluations included for newer models, so treat that comparison as directional.

This is the defensible meaning of “Gemini was dethroned”: both models arrived on the same date, and Muse moved two points ahead at the generally available xhigh setting and three at max on this specific independent index. It is not a claim that Muse wins every workload.

The counter-evidence sits on the same comparison page. Gemini 3.8 Flash high produces about 302 output tokens per second versus Muse’s 182, and its measured time to first token is 13.30 seconds versus 27.51 seconds. Gemini’s list rate is also lower at $0.75 per million fresh input tokens and $3.75 per million output tokens during its introductory period, compared with Muse standard at $1.25 and $4.25. Direct comparison

Four denominators
Each model wins a different line

Dots sit on separate scales; they are not connected because these metrics cannot be collapsed into one honest number.

Intelligence Index · 50–65 · higher →

DeepSeek 52
Gemini 59
Muse xh 61
Muse max 62
5065

Output speed · 100–320 tokens/s · higher →

DeepSeek 137
Muse 182
Gemini 302
100320

Time to first token · 0–30s · lower ←

DeepSeek 1.50
Gemini 13.30
Muse 27.51
0s30s

Pricing: Muse is efficient per task, not cheapest per token

Meta’s pricing page splits Muse into two commercial choices. Standard pricing excludes prompts and completions from model training. Contributor pricing is dramatically lower because the customer gives Meta permission to use that data to train future models. That is a data-governance choice, not a coupon code.

RouteCached input / 1MFresh input / 1MOutput / 1MKey condition
Muse Spark 1.3 standard$0.15$1.25$4.25Prompts/completions not used to train Meta models
Muse Spark 1.3 Contributor$0.002$0.10$0.20Training permission; lower rate limits
Gemini 3.8 Flash introductory$0.075$0.75$3.75Introductory through 31 Dec 2026
DeepSeek V4 Flash off-peak$0.007$0.22$0.66Outside specified weekday UTC peak windows
DeepSeek V4 Flash peak$0.014$0.44$1.3201:00–04:00 and 06:00–10:00 UTC, Mon–Fri

Meta rate card · Google rate card · DeepSeek rate card

For one illustrative long agent run with 1 million fresh input tokens and 200,000 output tokens, the token-only totals are $2.10 on Muse standard, $1.50 on Gemini, $0.704 on DeepSeek peak, $0.352 on DeepSeek off-peak, and $0.14 on Muse Contributor. The last number is commercially dramatic but unsuitable for confidential repositories or private customer data unless the organization has explicitly approved the training permission.

Illustrative workload
1M fresh input + 200K output

USD list price only. Shorter bars cost less. Excludes tools, retries, taxes, provider markup and human review.

Muse standard$2.10
Gemini intro$1.50
DeepSeek peak$0.704
DeepSeek off-peak$0.352
Muse Contributor$0.14

At a rounded ₹95/$ illustration: ₹199.50, ₹142.50, ₹66.88, ₹33.44 and ₹13.30 respectively, before GST and payment/FX fees.

Artificial Analysis still finds Muse xhigh slightly cheaper per completed index task: $0.55 versus Gemini high at $0.58. That is possible because completed-task cost depends on token consumption as well as the rate card. DeepSeek remains far lower at $0.11 per index task in the available comparison, although its current row lacks three evaluations used for the newer models. The right procurement formula is:

accepted-artifact cost = tokens + tools + retries + latency + human verification

Terminal-Bench 2.1 is useful. Terminal-Bench 4.0 is not “broken.”

The most confusing launch-day comparison is Gemini 3.8 Flash: Google reports 89.4% on Terminal-Bench 2.1 and 19.1% on Terminal-Bench 4.0. Claude Opus 5 moves from 89.1% to 51.8% in the same published table. The numbers look incompatible because the benchmark name hides a major change.

Version trap
89.4% and 19.1% are different measurements
89.4%
Terminal-Bench 2.1
89 tasks · Google self-computed · Terminus 2 default agent
19.1%
Terminal-Bench 4.0
66 changed tasks · official public leaderboard · different run provenance
Do not subtract these rows. The 70.3-point gap is not a measured regression in one fixed test.

Why 2.1 is good: it covers 89 real terminal tasks, runs agents in isolated environments, and uses executable verifiers to inspect the final container state. A confident narrative gets zero if the end state is wrong. It is public, repeatable, widely used, and one of the highest-weight components in Artificial Analysis’s current index. Terminal-Bench paper

Why 2.1 is no longer enough: a public benchmark gets easier to optimize against over time. Scores compress, tasks saturate, harness builders learn the failure modes, and data leakage becomes harder to rule out. A strong 2.1 score is still evidence of terminal competence; it is weaker evidence of unseen-task generalization.

What 4.0 fixes: Terminal-Bench’s maintainers describe it as a continuous benchmark. Version 3.0 replaced the old task set. Version 4.0 then fixed 19 tasks, removed eight—including saturated tasks and tasks with public solutions—and set a flat eight-hour timeout to reduce infrastructure noise. The current set contains 66 tasks. That maintenance makes the benchmark harder to game and the environment more stable. Maintainer explanation · repository

Why 4.0 can “not work well” for a launch comparison: it measures the model inside an agent, not the naked API. With only 66 expensive, long-horizon tasks, agent implementation, tool policy, refusals, output-token limits, resource configuration, and retry behavior can dominate. Google’s evaluation document says the 2.1 Gemini score was self-computed in Terminus 2, while its 4.0 score was copied from the official leaderboard’s best listed thinking level. That is not a matched A/B run. Google methodology

Muse has no official 4.0 result in Meta’s launch table. Until the same Muse and Gemini API checkpoints run through the same agent, tools, time budget, task release, and number of trials, Terminal-Bench 4.0 should answer “how did this complete agent stack perform?”—not “which base model is smartest?”

What X adds—and what it cannot prove

Mark Zuckerberg framed the release as a major coding and agentic jump, pointed readers to Muse Code and the API, and said Muse open weights are next. That is useful roadmap evidence, but still an executive launch claim.

Official launch post
Mark Zuckerberg on X · 2 Sep 2026

Artificial Analysis’s launch thread is more valuable for the purchasing decision because it puts the 62/61 index result beside task-level cost and identifies max as a limited preview.

Independent measurement
Artificial Analysis on X · 3 Sep 2026

Field posts provide a different kind of signal. In one public video comparison, @OmedVibeCodes criticized both Muse and Gemini and said Gemini required a second attempt to produce a working scene. The post is worth seeing because it challenges launch-day certainty. It is not a benchmark: there is no public repository, controlled prompt log, fixed time/token budget, or artifact-level rubric.

Launch-day field anecdote
@OmedVibeCodes on X · 3 Sep 2026

The useful pattern is not “X says the model is good” or “X says it failed.” It is that agentic systems need inspectable artifacts. A video can reveal a broken interaction or awkward design. It cannot quantify reliability without the prompt, code, logs, tests, retries, and acceptance rule.

Who should care

WorkloadBest starting candidateReason
Fast multimodal chat or interactive UIGemini 3.8 FlashHigher measured output speed, lower TTFT and broader audio support
Difficult agentic work where task success dominatesMuse Spark 1.3 xhighNarrow index lead and strong cost per completed index task
Low-cost API automation with non-peak schedulingDeepSeek V4 FlashFar lower list rate and 1M context
Private local or controlled deploymentDeepSeek V4 FlashMIT-licensed open weights; substantial hardware requirement
Disposable public-data experimentsMuse ContributorVery low rate if training permission is acceptable
Confidential code or customer dataMuse standard, Gemini paid, or governed DeepSeek routeAvoid Contributor until legal/security owners approve training use

What this changes

Muse Spark 1.3 makes Meta credible in the fast agentic-model tier again. Its xhigh route clears Gemini 3.8 Flash on the current Artificial Analysis index and nearly matches Gemini’s cost per completed index task despite higher list prices. Meta has also made a strategically aggressive bet with Contributor pricing.

But the release does not produce one obvious winner. Gemini still owns the speed line and the lower standard rate. DeepSeek owns the normal price line and open-weight control. Muse owns a narrow independent intelligence lead and an unusually cheap training-eligible route. The correct result is routing, not coronation.

Verdict: test the stack, not the screenshot

Put Muse Spark 1.3 on your shortlist if your workload is long-horizon coding, browsing, document production, or multi-step tool use. Start with xhigh, because that is the available configuration with an independent 61-point result. Do not budget around max until Meta publishes access and pricing.

Run the same 20–50 representative tasks through Muse, Gemini, and DeepSeek with one agent harness. Record task pass rate, output and cached tokens, wall-clock time, retries, refusal rate, human corrections, and accepted-artifact cost. Keep consequential actions behind tests, diffs, screenshots, or human approval.

Limitations: EyesTech did not run the three APIs for this analysis. Meta’s launch results are vendor-reported; Artificial Analysis uses its own harness; the DeepSeek comparison lacks three current index evaluations; X examples are anecdotal; and prices can change. Recheck max-mode access, open weights, pricing, and a matched Terminal-Bench 4.0 run at 7, 30, and 90 days.

Updated 3 September 2026. Confidence: High on release, price, and cited measurements; Medium on cross-harness purchasing implications.

Get the EyesTech Signal: model launches move fast. We track the version, harness, price, and deployment constraint that actually changes the decision.

Related EyesTech analysis: Claude Fable 5.1 · DeepSeek V4 vs Kimi K2.6

Categorized in:

A.I, Technology,

Last Update: September 3, 2026