Models / Muse Spark 1.1

Meta

Muse Spark 1.1

14 published results from 2 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Meta
Sources
2
Our benchmarks
0
Price
Not yet published

Reported by others

-0.02
Confirmed task success · rank 23 of 46
Unit
IPS effect estimate, higher is better
Range
-0.03 to -0.01
Sample
91237 observations
Configuration
Muse Spark 1.1
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.11
Praise over complaint · rank 35 of 46
Unit
IPS effect estimate, higher is better
Range
-0.12 to -0.10
Sample
38389 observations
Configuration
Muse Spark 1.1
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.06
Steerability · rank 33 of 46
Unit
IPS effect estimate, higher is better
Range
-0.07 to -0.05
Sample
124514 observations
Configuration
Muse Spark 1.1
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
6174897 observations
Configuration
Muse Spark 1.1
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,479
Overall · rank 3 of 44
Unit
Arena rating, higher is better
Range
1,471 to 1,488
Sample
4885 votes
Configuration
Muse Spark 1.1
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,463
Overall · rank 25 of 177
Unit
Arena rating, higher is better
Range
1,459 to 1,467
Sample
34415 votes
Configuration
Muse Spark 1.1
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,493
Business, management and finance · rank 2 of 402
Unit
Arena rating, higher is better
Range
1,485 to 1,501
Sample
6829 votes
Configuration
Muse Spark 1.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,449
Creative writing · rank 20 of 407
Unit
Arena rating, higher is better
Range
1,441 to 1,457
Sample
6925 votes
Configuration
Muse Spark 1.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,504
Expert prompts · rank 13 of 359
Unit
Arena rating, higher is better
Range
1,495 to 1,514
Sample
4099 votes
Configuration
Muse Spark 1.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,473
Instruction following · rank 13 of 409
Unit
Arena rating, higher is better
Range
1,467 to 1,479
Sample
12589 votes
Configuration
Muse Spark 1.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,491
Overall · rank 5 of 409
Unit
Arena rating, higher is better
Range
1,486 to 1,495
Sample
34760 votes
Configuration
Muse Spark 1.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,460
Writing, literature and language · rank 16 of 408
Unit
Arena rating, higher is better
Range
1,453 to 1,467
Sample
9304 votes
Configuration
Muse Spark 1.1
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by tau2-bench
20.6%
Consistency (pass^4) · rank 10 of 21
Unit
% of tasks, higher is better
Configuration
Muse Spark 1.1 (xhigh reasoning, tau2)
Measured
23 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
40.5%
Task success (pass^1) · rank 6 of 21
Unit
% of tasks, higher is better
Configuration
Muse Spark 1.1 (xhigh reasoning, tau2)
Measured
23 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.

Compare Muse Spark 1.1 with