Muse Spark 1.3 Benchmarks, Specifications & Availability
Explore Muse Spark 1.3 from Meta: published specifications, source-linked vendor benchmarks, independent evaluator coverage and recorded pricing when available.
Compare Muse Spark 1.3 with other models →Explore data coverage
Published specifications
- Provider
- Meta
- Access
- Proprietary
- License
- Proprietary
- Context window
- 1M
- Total parameters
- Not published
- Active parameters
- Not published
- Released
- 2026-09-02
- Modalities
- text, image, video, audio, file
- Family
- Muse Spark
Model card · Announcement · Website · OpenRouter
Model notes
Agentic/coding focused; Muse Code + Meta Model API. Contributor pricing tier on OpenRouter.
Pricing · OpenRouter
Dated cached OpenRouter rates in USD per 1M tokens. Open the dashboard for live enhancements. Per-metric endpoint minima can refer to different providers; they are not a guaranteed combined rate from one endpoint.
| Tier | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Date & source |
|---|---|---|---|---|---|
| Contributor (preferred) | [object Object] | [object Object] | [object Object] | — | 2026-10-03 · OpenRouter source |
| Default | [object Object] | [object Object] | [object Object] | — | 2026-10-03 · OpenRouter source |
Recorded pricing notes
min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03
OpenRouter contributor / BYOK-style tier slug.; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-01; min-healthy endpoint minima 2026-10-02; min-healthy endpoint minima 2026-10-03
Official / vendor benchmarks
Default headline records. Own-vendor, peer-vendor and third-party provenance remain visible in evidence; configurations may differ.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Agentic IF Index | 57.8%unknown effort | All 1 recorded result & sources57.8% · raw 57.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-02 Agentic IF Index; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy). https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| AutomationBench v1.0.6 | 49.6%max effort | All 1 recorded result & sources49.6% · raw 49.6 % Headline · max effort · Own vendor Source/record date: 2026-09-02 from chart/figure Meta Muse Spark 1.3 scorecard; AutomationBench E2E; max effort https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| DeepSearchQA | 90.3%max effort | All 1 recorded result & sources90.3% · raw 90.3 % Headline · max effort · Own vendor Source/record date: 2026-09-02 from chart/figure Meta Muse Spark 1.3 benchmark scorecard; max effort https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| DeepSWE v1.1 | 75.4%max effort | All 2 recorded results & sources75.4% · raw 75.4 % Headline · max effort · Own vendor Source/record date: 2026-09-02 Meta published table as transcribed by ExplainX; max reasoning; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 https://research.meta.ai/blog/introducing-muse-spark-1-375.4% · raw 75.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| GDPval-AA v2 | 1754unknown effort | All 1 recorded result & sources1754 · raw 1754 Elo Headline · unknown effort · Own vendor Source/record date: 2026-09-02 GDPVal-AA v2 Elo; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy). https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| JobBench | 64.9%unknown effort | All 1 recorded result & sources64.9% · raw 64.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-02 JobBench; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy). https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| MRCR 256K–512K | 98.5%unknown effort | All 1 recorded result & sources98.5% · raw 98.5 % Headline · unknown effort · Own vendor Source/record date: 2026-09-02 MRCR 256K–512K; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy). https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| MRCR 512K–1M | 98.1%unknown effort | All 1 recorded result & sources98.1% · raw 98.1 % Headline · unknown effort · Own vendor Source/record date: 2026-09-02 MRCR 512K–1M; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy). https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| Terminal-Bench 2.1 | 88.8%unknown effort | All 2 recorded results & sources88.8% · raw 88.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-02 Meta published table; ties GPT-5.6 Sol; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 Numbers transcribed from Meta announcement charts via ExplainX secondary writeup (Meta page is chart-heavy). https://research.meta.ai/blog/introducing-muse-spark-1-388.8% · raw 88.8 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI; matches the Meta figure value. https://naive.ai/en/research/ |
| OSWorld 2.0 | 66.9%max effort | All 1 recorded result & sources66.9% · raw 66.9 % Headline · max effort · Own vendor Source/record date: Not recorded from chart/figure Meta Muse Spark 1.3 scorecard; OSWorld 2.0 partial; max effort https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| OSWorld 2.0 (binary) | 32%max effort | All 1 recorded result & sources32% · raw 32 % Headline · max effort · Own vendor Source/record date: Not recorded from chart/figure Meta Muse Spark 1.3 scorecard; OSWorld 2.0 binary; max effort https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| SWE-Atlas QnA | 59.4%max effort | All 1 recorded result & sources59.4% · raw 59.4 % Headline · max effort · Own vendor Source/record date: Not recorded from chart/figure Meta Muse Spark 1.3 scorecard; SWEAtlas CodeBase QnA; max effort https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| AA Intelligence Index (vendor-cited) | 62unknown effort | All 1 recorded result & sources62 · raw 62 index Headline · unknown effort · Peer vendor Source/record date: Not recorded from chart/figure as reported on Ling-3.0-flash-VL HF card AA Index v4.1.1 chart (peer spillover); Muse Spark 1.3 (max) https://huggingface.co/inclusionAI/Ling-3.0-flash-VL |
| AutomationBench | 49.6%max effort | All 1 recorded result & sources49.6% · raw 49.6 % Headline · max effort · Own vendor Source/record date: 2026-09-02 Vendor scorecard figure (from chart/figure); max effort. Same figure: Muse Spark 1.2 38.2, GPT-5.6 Sol 46.7, Opus 5 50.3 (peers not in table). https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| Agents' Last Exam | 33.3%unknown effort | All 1 recorded result & sources33.3% · raw 33.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| Gray Swan IPI | 15.9%unknown effort | All 1 recorded result & sources15.9% · raw 15.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon Gray Swan chart (K=15 attack success rate) https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini_4_cyber_evals_gray_swan_i.width-1200.format-webp.webp |
Independent evaluators
Evaluator harnesses are distinct from vendor measurements. Missing coverage is not a failed test.
| Benchmark / evaluator | Headline score | Evidence |
|---|---|---|
| Intelligence Index · Artificial Analysis | 48max effort | All 2 recorded results & sources48 · raw 48 index Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/muse-spark-1-345 · raw 45 index Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/muse-spark-1-3-xhigh |
| Cost per Intelligence Index task · Artificial Analysis | $1.60max effort | All 2 recorded results & sources$1.60 · raw 1.6 USD Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/muse-spark-1-3$1.37 · raw 1.37 USD Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; https://artificialanalysis.ai/models/muse-spark-1-3-xhigh |
| Output speed · Artificial Analysis | 152 tok/smax effort | All 2 recorded results & sources152 tok/s · raw 152 tok/s Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/muse-spark-1-3140 tok/s · raw 140 tok/s Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Current AA model leaderboard; Median output tokens/s; leaderboard rounds to whole tokens. https://artificialanalysis.ai/models/muse-spark-1-3-xhigh |
| Vals Index · Vals AI | 58.2%max effort | All 2 recorded results & sources58.2% · raw 58.16 % Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Max variant. Cost/test $3.79. https://www.vals.ai/benchmarks/vals_index53.2% · raw 53.2 % Alternative · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Cost/test $3.38. https://www.vals.ai/benchmarks/vals_index |
| Bugs fixed /105 · Bug Hunt Bench | 32.2 fixesmax effort | All 5 recorded results & sources32.2 fixes · raw 32.2 fixes Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Harness: Muse Code / Meta API; effort max; 5 runs; evaluation 2026-09-17. Best documented score for this effort in Oct 1 README. Headline: best documented model run. https://github.com/phuryn/bug-hunt-bench20.3 fixes · raw 20.3 fixes Alternative · xhigh effort · Independent evaluator Source/record date: 2026-10-03 Harness: Muse Code / Meta API; effort xhigh; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench18.7 fixes · raw 18.7 fixes Alternative · high effort · Independent evaluator Source/record date: 2026-10-03 Harness: Muse Code / Meta API; effort high; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench13 fixes · raw 13 fixes Alternative · medium effort · Independent evaluator Source/record date: 2026-10-03 Harness: Muse Code / Meta API; effort medium; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench9.7 fixes · raw 9.7 fixes Alternative · low effort · Independent evaluator Source/record date: 2026-10-03 Harness: Muse Code / Meta API; effort low; 3 runs; evaluation 2026-09-14. Best documented score for this effort in Oct 1 README. https://github.com/phuryn/bug-hunt-bench |
| Vibe Code Bench v1.1 · Vals AI | 85.9%max effort | All 2 recorded results & sources85.9% · raw 85.86 % Headline · max effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Harness: OpenHands. Max variant. Cost/test $2.54. https://www.vals.ai/benchmarks/vibe-code82.9% · raw 82.86 % Alternative · unknown effort · Independent evaluator Source/record date: 2026-10-03 Refreshed from current Vals leaderboard. Harness: OpenHands. Cost/test $2.10. https://www.vals.ai/benchmarks/vibe-code |
| CUDA board % of roofline · KernelBench (community board) | 4.9%unknown effort | All 1 recorded result & sources4.9% · raw 4.9 % Headline · unknown effort · Independent evaluator Source/record date: 2026-09-23 CUDA 3/4 muse·ultra; Mega 2.18× https://kernelbench.com/models/muse-spark-1.3 |
| GDPval-AA Elo · Artificial Analysis | 1674max effort | All 1 recorded result & sources1674 · raw 1674 Elo Headline · max effort · Independent evaluator Source/record date: 2026-09-23 GDPval-AA v2.1 Elo; Muse Spark 1.3 (max) https://artificialanalysis.ai/evaluations/gdpval-aa |
Read how we select and source scores or the comparison guide.