Models / Mistral Large 4

Mistral

Mistral Large 4

104 published results from 4 sources. Each card shows where the number comes from and what it does not measure. The overall leaderboard combines them; here each stands alone.

Provider
Mistral
Sources
4
Our benchmarks
4
Price per million tokens
Not listed on OpenRouter

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
$0.0082
Cost per game · rank 2 of 11
Unit
US dollars, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
0.0%
Games lost on time · rank 1 of 11
Unit
% of games, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.6%
Invalid moves · rank 10 of 11
Unit
% of moves, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
555
Ladder Elo · rank 8 of 11
Unit
ladder Elo, higher is better
Range
392 to 727
Sample
8 games
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.8 s
Median move time · rank 3 of 11
Unit
milliseconds, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · BulletBench
$0.0117
Cost per game · rank 12 of 16
Unit
US dollars, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
58.3%
Games lost on time · rank 11 of 16
Unit
% of games, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
1.8%
Invalid moves · rank 14 of 16
Unit
% of moves, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
249
Ladder Elo · rank 11 of 16
Unit
ladder Elo, higher is better
Range
0 to 486
Sample
12 games
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.8 s
Median move time · rank 8 of 16
Unit
milliseconds, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
1.2%
Channel rules broken · rank 14 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
4.2%
Channel rules broken · rank 19 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
31.6%
Claims to check · rank 19 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.3%
Claims to check · rank 13 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
95.9%
Content quality · rank 14 of 20
Unit
% of checks, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
89.2%
Content quality · rank 19 of 20
Unit
% of checks, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
7.7%
Failed outputs · rank 19 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
0.0%
Missing UK information · rank 1 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Missing UK information · rank 1 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
6.0%
Not findable · rank 13 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.0%
Not findable · rank 6 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
2.4%
Publish-ready listings · rank 19 of 20
Unit
% of products, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
10.7%
Publish-ready listings · rank 15 of 20
Unit
% of products, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Reliably publish-ready · rank 17 of 20
Unit
% of products, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.8%
Reliably publish-ready · rank 15 of 20
Unit
% of products, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
97.0%
Unsupported claims · rank 20 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
80.4%
Unsupported claims · rank 14 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
6.74
Unsupported claims · rank 20 of 20
Unit
claims per product, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.69
Unsupported claims · rank 16 of 20
Unit
claims per product, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
42.3%
Wrong attributes · rank 19 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
17.3%
Wrong attributes · rank 9 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.8%
Wrong category or variant · rank 3 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
2.4%
Wrong category or variant · rank 9 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 12 of 168 attempts; invalid JSON: 1 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Channel compliance · rank 1 of 20
Unit
% of products, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
98.0%
Channel compliance · rank 20 of 20
Unit
% of products, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.0%
Channel rules broken · rank 18 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
6.5%
Channel rules broken · rank 20 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
12.7%
Claims to check · rank 19 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.7%
Claims to check · rank 14 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 20
Unit
% of conflicts, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 20
Unit
% of conflicts, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
96.5%
Content quality · rank 12 of 20
Unit
% of checks, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
90.6%
Content quality · rank 19 of 20
Unit
% of checks, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0033
Cost per product · rank 3 of 20
Unit
US dollars, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
$0.0125
Cost per product · rank 11 of 20
Unit
US dollars, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
93.1%
Decision accuracy · rank 16 of 20
Unit
% of decisions, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
88.7%
Decision accuracy · rank 19 of 20
Unit
% of decisions, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
7.1%
Failed outputs · rank 19 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
94.4%
Field accuracy · rank 3 of 20
Unit
% of missing fields, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
86.9%
Field accuracy · rank 19 of 20
Unit
% of missing fields, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.5%
Invented values · rank 16 of 20
Unit
% of filled values, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
4.4%
Invented values · rank 18 of 20
Unit
% of filled values, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Missing UK information · rank 1 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Missing UK information · rank 1 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
4.2%
Not findable · rank 6 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
5.4%
Not findable · rank 9 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
18.4%
Publish-ready listings · rank 18 of 20
Unit
% of products, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
39.9%
Publish-ready listings · rank 15 of 20
Unit
% of products, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
7.1%
Reliably publish-ready · rank 18 of 20
Unit
% of products, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
10.7%
Reliably publish-ready · rank 17 of 20
Unit
% of products, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
64.3%
Unsupported claims · rank 19 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
33.3%
Unsupported claims · rank 13 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
2.07
Unsupported claims · rank 18 of 20
Unit
claims per product, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.91
Unsupported claims · rank 12 of 20
Unit
claims per product, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
38.1%
Wrong attributes · rank 19 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
19.6%
Wrong attributes · rank 13 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.2%
Wrong category or variant · rank 1 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
4.2%
Wrong category or variant · rank 16 of 20
Unit
% of products, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Failures
reply cut off at the token limit: 6 of 168 attempts; provider error after retries: 4 of 168 attempts; invalid JSON: 2 of 168 attempts
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · DeckBench
0.0%
Accurate decks · rank 15 of 18
Unit
% of tasks, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
16.7%
Caveat dropped · rank 10 of 18
Unit
% of tasks, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
33.3%
Clean layout · rank 12 of 18
Unit
% of tasks, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
$0.0148
Cost per deck · rank 3 of 18
Unit
US dollars, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
515
Deck rating · rank 18 of 18
Unit
rating, higher is better
Range
296 to 615
Sample
84 comparisons
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
30.5%
Design quality · rank 18 of 18
Unit
% of the maximum, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
16.7%
Draft figure quoted · rank 8 of 18
Unit
% of tasks, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
0.0%
Findings missing · rank 1 of 18
Unit
% of tasks, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
10.7%
Head-to-head win rate · rank 18 of 18
Unit
% of comparisons, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
66.7%
Layout defects · rank 12 of 18
Unit
% of tasks, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
0.0%
Misleading metric used · rank 1 of 18
Unit
% of tasks, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
0.0%
Presentable decks · rank 10 of 18
Unit
% of tasks, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
83.3%
Recommendation late or wrong · rank 17 of 18
Unit
% of tasks, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
100.0%
Slides needing work · rank 18 of 18
Unit
% of tasks, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
50.0%
Unsupported claims · rank 16 of 18
Unit
% of tasks, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · DeckBench
33.3%
Unsupported numbers · rank 18 of 18
Unit
% of tasks, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
Measured by Spring Prompt · ROASBench
55.8
Audience score · rank 11 of 19
Unit
score out of 100, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
51.3
Audience score · rank 13 of 19
Unit
score out of 100, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
34.2
Business score · rank 10 of 19
Unit
score out of 100, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
19.3
Business score · rank 14 of 19
Unit
score out of 100, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
16.7
Consistency score · rank 14 of 19
Unit
score out of 100, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
9.7
Consistency score · rank 16 of 19
Unit
score out of 100, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
$533,401
Contribution profit · rank 9 of 19
Unit
simulated US dollars, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
$156,831
Contribution profit · rank 14 of 19
Unit
simulated US dollars, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
$0.0657
Cost of a run · rank 4 of 19
Unit
US dollars, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
$0.19
Cost of a run · rank 9 of 19
Unit
US dollars, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
3
Months over budget · rank 18 of 19
Unit
months of 12, lower is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
0
Months over budget · rank 1 of 19
Unit
months of 12, lower is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
37.0
Overall score · rank 10 of 19
Unit
score out of 100, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
27.4
Overall score · rank 14 of 19
Unit
score out of 100, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
54.4
Planning score · rank 17 of 19
Unit
score out of 100, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
55.0
Planning score · rank 3 of 19
Unit
score out of 100, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
$0.45
Return on ad spend · rank 9 of 19
Unit
profit per $1 spent, higher is better
Configuration
Mistral Large 4, provider default reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
$0.14
Return on ad spend · rank 14 of 19
Unit
profit per $1 spent, higher is better
Configuration
Mistral Large 4 (high reasoning), high reasoning
Measured
6 Oct 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Compare Mistral Large 4 with