Models / Muse Spark 1.3

Meta

Muse Spark 1.3

77 published results from 7 sources. Each card shows where the number comes from and what it does not measure. The overall leaderboard combines them; here each stands alone.

Provider
Meta
Sources
7
Our benchmarks
3
Price per million tokens
$1.25 in · $4.25 out
OpenRouter list price, 1 Oct 2026 · 1,048,576-token context

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
$0.0142
Cost per game · rank 10 of 24
Unit
US dollars, lower is better
Configuration
Muse Spark 1.3 (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
91.7%
Games lost on time · rank 22 of 24
Unit
% of games, lower is better
Configuration
Muse Spark 1.3 (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.0%
Invalid moves · rank 1 of 24
Unit
% of moves, lower is better
Configuration
Muse Spark 1.3 (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
33
Ladder Elo · rank 12 of 24
Unit
ladder Elo, higher is better
Range
0 to 270
Sample
12 games
Configuration
Muse Spark 1.3 (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
4.2 s
Median move time · rank 24 of 24
Unit
milliseconds, lower is better
Configuration
Muse Spark 1.3 (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
0.0%
Channel rules broken · rank 1 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Claims to check · rank 1 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
98.2%
Content quality · rank 4 of 18
Unit
% of checks, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
0.0%
Missing UK information · rank 1 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
5.4%
Not findable · rank 9 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
34.5%
Publish-ready listings · rank 7 of 18
Unit
% of products, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
10.7%
Reliably publish-ready · rank 9 of 18
Unit
% of products, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
57.7%
Unsupported claims · rank 8 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.50
Unsupported claims · rank 8 of 18
Unit
claims per product, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
14.3%
Wrong attributes · rank 4 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.6%
Wrong category or variant · rank 14 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Channel compliance · rank 1 of 18
Unit
% of products, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Channel rules broken · rank 1 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Claims to check · rank 1 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 18
Unit
% of conflicts, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
97.7%
Content quality · rank 6 of 18
Unit
% of checks, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0164
Cost per product · rank 10 of 18
Unit
US dollars, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
97.5%
Decision accuracy · rank 5 of 18
Unit
% of decisions, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
92.9%
Field accuracy · rank 10 of 18
Unit
% of missing fields, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.7%
Invented values · rank 5 of 18
Unit
% of filled values, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Missing UK information · rank 1 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
6.0%
Not findable · rank 10 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
65.5%
Publish-ready listings · rank 7 of 18
Unit
% of products, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
48.2%
Reliably publish-ready · rank 7 of 18
Unit
% of products, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
14.9%
Unsupported claims · rank 7 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.43
Unsupported claims · rank 5 of 18
Unit
claims per product, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
16.7%
Wrong attributes · rank 8 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
4.2%
Wrong category or variant · rank 15 of 18
Unit
% of products, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · ROASBench
51.6
Audience score · rank 11 of 17
Unit
score out of 100, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
25.8
Business score · rank 11 of 17
Unit
score out of 100, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
26.7
Consistency score · rank 11 of 17
Unit
score out of 100, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
$337,251
Contribution profit · rank 10 of 17
Unit
simulated US dollars, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
$0.18
Cost of a run · rank 7 of 17
Unit
US dollars, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
0
Months over budget · rank 1 of 17
Unit
months of 12, lower is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
34.1
Overall score · rank 11 of 17
Unit
score out of 100, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
54.8
Planning score · rank 6 of 17
Unit
score out of 100, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
$0.29
Return on ad spend · rank 10 of 17
Unit
profit per $1 spent, higher is better
Configuration
Muse Spark 1.3, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

Reported by APEX-Agents
57.8%
Tasks passed · rank 13 of 39
Unit
% of tasks, higher is better
Configuration
Muse Spark 1.3
Measured
1 Oct 2026
Not shown
Not your firm's documents or tools; graded by rubric, not by a client.
0.08
Confirmed task success · rank 4 of 46
Unit
IPS effect estimate, higher is better
Range
0.07 to 0.10
Sample
43358 observations
Configuration
Muse Spark 1.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.04
Praise over complaint · rank 10 of 46
Unit
IPS effect estimate, higher is better
Range
0.02 to 0.06
Sample
15832 observations
Configuration
Muse Spark 1.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.02
Steerability · rank 11 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.03
Sample
45532 observations
Configuration
Muse Spark 1.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
5211195 observations
Configuration
Muse Spark 1.3
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,468
Overall · rank 6 of 44
Unit
Arena rating, higher is better
Range
1,450 to 1,486
Sample
1006 votes
Configuration
Muse Spark 1.3 (max)
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,479
Overall · rank 4 of 177
Unit
Arena rating, higher is better
Range
1,473 to 1,485
Sample
9672 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,506
Business, management and finance · rank 1 of 402
Unit
Arena rating, higher is better
Range
1,492 to 1,519
Sample
1912 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,450
Creative writing · rank 13 of 407
Unit
Arena rating, higher is better
Range
1,436 to 1,464
Sample
2035 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,527
Expert prompts · rank 1 of 359
Unit
Arena rating, higher is better
Range
1,510 to 1,544
Sample
1149 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,486
Instruction following · rank 4 of 409
Unit
Arena rating, higher is better
Range
1,476 to 1,496
Sample
3726 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,494
Overall · rank 2 of 409
Unit
Arena rating, higher is better
Range
1,488 to 1,501
Sample
10036 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,466
Writing, literature and language · rank 9 of 408
Unit
Arena rating, higher is better
Range
1,454 to 1,478
Sample
2655 votes
Configuration
Muse Spark 1.3 (max)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
75.8
Artificial Analysis Coding Index · rank 25 of 151
Unit
index score, higher is better
Configuration
Muse Spark 1.3 (max)
Measured
1 Oct 2026
Not shown
Not your codebase or tools.
76.5
Artificial Analysis Coding Index · rank 15 of 151
Unit
index score, higher is better
Configuration
Muse Spark 1.3 (xhigh)
Measured
1 Oct 2026
Not shown
Not your codebase or tools.
48.1
Artificial Analysis Intelligence Index · rank 22 of 314
Unit
index score, higher is better
Configuration
Muse Spark 1.3 (max)
Measured
1 Oct 2026
Not shown
Not business work, and a blend: read the parts for any one task.
45.1
Artificial Analysis Intelligence Index · rank 33 of 314
Unit
index score, higher is better
Configuration
Muse Spark 1.3 (xhigh)
Measured
1 Oct 2026
Not shown
Not business work, and a blend: read the parts for any one task.
93.5%
GPQA Diamond · rank 14 of 276
Unit
% of questions, higher is better
Configuration
Muse Spark 1.3 (max)
Measured
1 Oct 2026
Not shown
Not applied work; multiple-choice science questions.
94.1%
GPQA Diamond · rank 7 of 276
Unit
% of questions, higher is better
Configuration
Muse Spark 1.3 (xhigh)
Measured
1 Oct 2026
Not shown
Not applied work; multiple-choice science questions.
48.7%
Humanity's Last Exam · rank 28 of 314
Unit
% of questions, higher is better
Configuration
Muse Spark 1.3 (max)
Measured
1 Oct 2026
Not shown
Not everyday work; academic questions at the edge of expertise.
47.5%
Humanity's Last Exam · rank 34 of 314
Unit
% of questions, higher is better
Configuration
Muse Spark 1.3 (xhigh)
Measured
1 Oct 2026
Not shown
Not everyday work; academic questions at the edge of expertise.
83.0%
Long-context reasoning (AA-LCR) · rank 21 of 309
Unit
% of questions, higher is better
Configuration
Muse Spark 1.3 (max)
Measured
1 Oct 2026
Not shown
Not retrieval over your own document store.
83.0%
Long-context reasoning (AA-LCR) · rank 21 of 309
Unit
% of questions, higher is better
Configuration
Muse Spark 1.3 (xhigh)
Measured
1 Oct 2026
Not shown
Not retrieval over your own document store.
58.8%
SciCode · rank 14 of 153
Unit
% of problems, higher is better
Configuration
Muse Spark 1.3 (max)
Measured
1 Oct 2026
Not shown
Not general software engineering.
59.7%
SciCode · rank 10 of 153
Unit
% of problems, higher is better
Configuration
Muse Spark 1.3 (xhigh)
Measured
1 Oct 2026
Not shown
Not general software engineering.
84.3%
Terminal-Bench 2.1 · rank 30 of 151
Unit
% of tasks, higher is better
Configuration
Muse Spark 1.3 (max)
Measured
1 Oct 2026
Not shown
Not other harnesses or tools; superseded by 4.0 for newer models.
85.4%
Terminal-Bench 2.1 · rank 25 of 151
Unit
% of tasks, higher is better
Configuration
Muse Spark 1.3 (xhigh)
Measured
1 Oct 2026
Not shown
Not other harnesses or tools; superseded by 4.0 for newer models.
33.3%
Terminal-Bench 4.0 · rank 34 of 150
Unit
% of tasks, higher is better
Configuration
Muse Spark 1.3 (max)
Measured
1 Oct 2026
Not shown
Not other harnesses or tools; one attempt per task.
16.7%
Terminal-Bench 4.0 · rank 56 of 150
Unit
% of tasks, higher is better
Configuration
Muse Spark 1.3 (xhigh)
Measured
1 Oct 2026
Not shown
Not other harnesses or tools; one attempt per task.
50.5%
Τ-bench banking · rank 3 of 142
Unit
% of tasks, higher is better
Configuration
Muse Spark 1.3 (max)
Measured
1 Oct 2026
Not shown
Not your policies or systems; a simulated customer.
47.2%
Τ-bench banking · rank 9 of 142
Unit
% of tasks, higher is better
Configuration
Muse Spark 1.3 (xhigh)
Measured
1 Oct 2026
Not shown
Not your policies or systems; a simulated customer.
Reported by GDP.pdf
27.6%
Professional document tasks · rank 7 of 49
Unit
% of rubric, higher is better
Configuration
Muse Spark 1.3
Measured
1 Oct 2026
Not shown
Not your documents; 100 tasks, so small differences are noise.
Reported by GDP.pdf
27.6%
Professional document tasks · rank 7 of 49
Unit
% of rubric, higher is better
Configuration
Muse Spark 1.3 (xhigh)
Measured
1 Oct 2026
Not shown
Not your documents; 100 tasks, so small differences are noise.

Compare Muse Spark 1.3 with