Examenos

Muse Spark 1.3 Benchmarks, Specifications & Availability

Explore Muse Spark 1.3 from Meta: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.

Compare Muse Spark 1.3 with other models →Explore data coverage

Published specifications

Provider
Meta
Access
Proprietary
License
Proprietary
Context window
1M
Total parameters
Not published
Active parameters
Not published
Released
2026-09-02
Modalities
text, image, video, audio, file
Family
Muse Spark

Model card · Announcement · Website · OpenRouter

Model notes

Agentic/coding focused; Muse Code + Meta Model API. Contributor pricing tier on OpenRouter.

Pricing · OpenRouter

Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.

Recorded pricing tiers
TierInput / 1MOutput / 1MCached input / 1MCache write / 1MDate & source
Contributor (preferred)[object Object][object Object][object Object]—2026-10-03 · OpenRouter source
Default[object Object][object Object][object Object]—2026-10-03 · OpenRouter source
Recorded pricing notes

min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03

OpenRouter contributor / BYOK-style tier slug.; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03

Official / vendor benchmarks

Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.

Official / vendor headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Agentic IF Index57.8%unknown effort
All 1 recorded result & sources

57.8% · raw 57.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-02

Agentic IF Index; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy).

https://research.meta.ai/blog/introducing-muse-spark-1-3
AutomationBench v1.0.649.6%max effort
All 1 recorded result & sources

49.6% · raw 49.6 %

Headline · max effort · Own vendor

Source/record date: 2026-09-02

from chart/figure Meta Muse Spark 1.3 scorecard; AutomationBench E2E; max effort

https://research.meta.ai/blog/introducing-muse-spark-1-3
DeepSearchQA90.3%max effort
All 1 recorded result & sources

90.3% · raw 90.3 %

Headline · max effort · Own vendor

Source/record date: 2026-09-02

from chart/figure Meta Muse Spark 1.3 benchmark scorecard; max effort

https://research.meta.ai/blog/introducing-muse-spark-1-3
DeepSWE v1.175.4%max effort
All 2 recorded results & sources

75.4% · raw 75.4 %

Headline · max effort · Own vendor

Source/record date: 2026-09-02

Meta published table as transcribed by ExplainX; max reasoning; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3

https://research.meta.ai/blog/introducing-muse-spark-1-3

75.4% · raw 75.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
GDPval-AA v21754unknown effort
All 1 recorded result & sources

1754 · raw 1754 Elo

Headline · unknown effort · Own vendor

Source/record date: 2026-09-02

GDPVal-AA v2 Elo; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy).

https://research.meta.ai/blog/introducing-muse-spark-1-3
JobBench64.9%unknown effort
All 1 recorded result & sources

64.9% · raw 64.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-02

JobBench; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy).

https://research.meta.ai/blog/introducing-muse-spark-1-3
MRCR 256K–512K98.5%unknown effort
All 1 recorded result & sources

98.5% · raw 98.5 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-02

MRCR 256K–512K; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy).

https://research.meta.ai/blog/introducing-muse-spark-1-3
MRCR 512K–1M98.1%unknown effort
All 1 recorded result & sources

98.1% · raw 98.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-02

MRCR 512K–1M; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy).

https://research.meta.ai/blog/introducing-muse-spark-1-3
Terminal-Bench 2.188.8%unknown effort
All 2 recorded results & sources

88.8% · raw 88.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-02

Meta published table; ties GPT-5.6 Sol; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy).

https://research.meta.ai/blog/introducing-muse-spark-1-3

88.8% · raw 88.8 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI; matches the Meta figure value.

https://naive.ai/en/research/
OSWorld 2.066.9%max effort
All 1 recorded result & sources

66.9% · raw 66.9 %

Headline · max effort · Own vendor

Source/record date: Not recorded

from chart/figure Meta Muse Spark 1.3 scorecard; OSWorld 2.0 partial; max effort

https://research.meta.ai/blog/introducing-muse-spark-1-3
OSWorld 2.0 (binary)32%max effort
All 1 recorded result & sources

32% · raw 32 %

Headline · max effort · Own vendor

Source/record date: Not recorded

from chart/figure Meta Muse Spark 1.3 scorecard; OSWorld 2.0 binary; max effort

https://research.meta.ai/blog/introducing-muse-spark-1-3
SWE-Atlas QnA59.4%max effort
All 1 recorded result & sources

59.4% · raw 59.4 %

Headline · max effort · Own vendor

Source/record date: Not recorded

from chart/figure Meta Muse Spark 1.3 scorecard; SWEAtlas CodeBase QnA; max effort

https://research.meta.ai/blog/introducing-muse-spark-1-3
AA Intelligence Index (vendor-cited)62unknown effort
All 1 recorded result & sources

62 · raw 62 index

Headline · unknown effort · Peer vendor

Source/record date: Not recorded

from chart/figure as reported on Ling-3.0-flash-VL HF card AA Index v4.1.1 chart (peer spillover); Muse Spark 1.3 (max)

https://huggingface.co/inclusionAI/Ling-3.0-flash-VL
AutomationBench49.6%max effort
All 1 recorded result & sources

49.6% · raw 49.6 %

Headline · max effort · Own vendor

Source/record date: 2026-09-02

Vendor scorecard figure (from chart/figure); max effort. Same figure: Muse Spark 1.2 38.2, GPT-5.6 Sol 46.7, Opus 5 50.3 (peers not in table).

https://research.meta.ai/blog/introducing-muse-spark-1-3
Agents' Last Exam33.3%unknown effort
All 1 recorded result & sources

33.3% · raw 33.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
Gray Swan IPI15.9%unknown effort
All 1 recorded result & sources

15.9% · raw 15.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon Gray Swan chart (K=15 attack success rate)

https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp

Independent evaluators

Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.

Independent evaluator headline scores; expand evidence for every record
Benchmark / evaluatorHeadline scoreEvidence
Intelligence Index · Artificial Analysis48max effort
All 2 recorded results & sources

48 · raw 48 index

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/muse-spark-1-3

45 · raw 45 index

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/muse-spark-1-3-xhigh
Cost per Intelligence Index task · Artificial Analysis$1.60max effort
All 2 recorded results & sources

$1.60 · raw 1.6 USD

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/muse-spark-1-3

$1.37 · raw 1.37 USD

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard;

https://artificialanalysis.ai/models/muse-spark-1-3-xhigh
Output speed · Artificial Analysis152 tok/smax effort
All 2 recorded results & sources

152 tok/s · raw 152 tok/s

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/muse-spark-1-3

140 tok/s · raw 140 tok/s

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens.

https://artificialanalysis.ai/models/muse-spark-1-3-xhigh
Vals Index · Vals AI58.2%max effort
All 2 recorded results & sources

58.2% · raw 58.16 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Max variant. Cost/test $3.79.

https://www.vals.ai/benchmarks/vals_index

53.2% · raw 53.2 %

Alternative · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Cost/test $3.38.

https://www.vals.ai/benchmarks/vals_index
Bugs fixed /105 · Bug Hunt Bench32.2 fixesmax effort
All 5 recorded results & sources

32.2 fixes · raw 32.2 fixes

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Muse Code / Meta API; effort max; 5 runs; evaluation 2026-09-17. Best documented score for this effort in Oct 1 README. Headline: best documented model run.

https://github.com/phuryn/bug-hunt-bench

20.3 fixes · raw 20.3 fixes

Alternative · xhigh effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Muse Code / Meta API; effort xhigh; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

18.7 fixes · raw 18.7 fixes

Alternative · high effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Muse Code / Meta API; effort high; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

13 fixes · raw 13 fixes

Alternative · medium effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Muse Code / Meta API; effort medium; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench

9.7 fixes · raw 9.7 fixes

Alternative · low effort · Independent evaluator

Source/record date: 2026-10-03

Harness: Muse Code / Meta API; effort low; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README.

https://github.com/phuryn/bug-hunt-bench
Vibe Code Bench v1.1 · Vals AI85.9%max effort
All 2 recorded results & sources

85.9% · raw 85.86 %

Headline · max effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Max variant. Cost/test $2.54.

https://www.vals.ai/benchmarks/vibe-code

82.9% · raw 82.86 %

Alternative · unknown effort · Independent evaluator

Source/record date: 2026-10-03

Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $2.10.

https://www.vals.ai/benchmarks/vibe-code
CUDA board % of roofline · KernelBench (community board)4.9%unknown effort
All 1 recorded result & sources

4.9% · raw 4.9 %

Headline · unknown effort · Independent evaluator

Source/record date: 2026-09-23

CUDA 3/4 muse·ultra; Mega 2.18×

https://kernelbench.com/models/muse-spark-1.3
GDPval-AA Elo · Artificial Analysis1674max effort
All 1 recorded result & sources

1674 · raw 1674 Elo

Headline · max effort · Independent evaluator

Source/record date: 2026-09-23

GDPval-AA v2.1 Elo; Muse Spark 1.3 (max)

https://artificialanalysis.ai/evaluations/gdpval-aa

Read how we select and source scores or the comparison guide.