Models / Mistral Medium 3.5

Mistral

Mistral Medium 3.5

93 published results from 6 sources. Each card shows where the number comes from and what it does not measure. The overall leaderboard combines them; here each stands alone.

Provider
Mistral
Sources
6
Our benchmarks
3
Price per million tokens
$1.50 in · $7.50 out
OpenRouter list price, 1 Oct 2026 · 262,144-token context

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
$0.0261
Cost per game · rank 3 of 10
Unit
US dollars, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
0.0%
Games lost on time · rank 1 of 10
Unit
% of games, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.4%
Invalid moves · rank 8 of 10
Unit
% of moves, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
555
Ladder Elo · rank 2 of 10
Unit
ladder Elo, higher is better
Range
396 to 765
Sample
8 games
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.5 s
Median move time · rank 1 of 10
Unit
milliseconds, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · BulletBench
$0.0420
Cost per game · rank 20 of 24
Unit
US dollars, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
$0.0354
Cost per game · rank 19 of 24
Unit
US dollars, lower is better
Configuration
Mistral Medium 3.5 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
8.3%
Games lost on time · rank 6 of 24
Unit
% of games, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.0%
Games lost on time · rank 1 of 24
Unit
% of games, lower is better
Configuration
Mistral Medium 3.5 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.8%
Invalid moves · rank 22 of 24
Unit
% of moves, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.4%
Invalid moves · rank 19 of 24
Unit
% of moves, lower is better
Configuration
Mistral Medium 3.5 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
385
Ladder Elo · rank 5 of 24
Unit
ladder Elo, higher is better
Range
276 to 464
Sample
12 games
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
572
Ladder Elo · rank 2 of 24
Unit
ladder Elo, higher is better
Range
428 to 708
Sample
12 games
Configuration
Mistral Medium 3.5 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.5 s
Median move time · rank 4 of 24
Unit
milliseconds, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · BulletBench
0.5 s
Median move time · rank 3 of 24
Unit
milliseconds, lower is better
Configuration
Mistral Medium 3.5 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · BulletBench
$0.0297
Cost per game · rank 14 of 15
Unit
US dollars, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
$0.0491
Cost per game · rank 15 of 15
Unit
US dollars, lower is better
Configuration
Mistral Medium 3.5 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
0.0%
Games lost on time · rank 1 of 15
Unit
% of games, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.0%
Games lost on time · rank 1 of 15
Unit
% of games, lower is better
Configuration
Mistral Medium 3.5 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.3%
Invalid moves · rank 10 of 15
Unit
% of moves, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.0%
Invalid moves · rank 1 of 15
Unit
% of moves, lower is better
Configuration
Mistral Medium 3.5 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
529
Ladder Elo · rank 1 of 15
Unit
ladder Elo, higher is better
Range
308 to 680
Sample
12 games
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
529
Ladder Elo · rank 1 of 15
Unit
ladder Elo, higher is better
Range
361 to 669
Sample
12 games
Configuration
Mistral Medium 3.5 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.5 s
Median move time · rank 3 of 15
Unit
milliseconds, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · BulletBench
0.5 s
Median move time · rank 4 of 15
Unit
milliseconds, lower is better
Configuration
Mistral Medium 3.5 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
1.2%
Channel rules broken · rank 14 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
25.6%
Claims to check · rank 16 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
94.1%
Content quality · rank 15 of 18
Unit
% of checks, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
0.0%
Missing UK information · rank 1 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
10.1%
Not findable · rank 16 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
17.3%
Publish-ready listings · rank 11 of 18
Unit
% of products, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
10.7%
Reliably publish-ready · rank 9 of 18
Unit
% of products, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
76.8%
Unsupported claims · rank 12 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.55
Unsupported claims · rank 14 of 18
Unit
claims per product, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
36.3%
Wrong attributes · rank 17 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
13.1%
Wrong category or variant · rank 18 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Channel compliance · rank 1 of 18
Unit
% of products, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Channel rules broken · rank 1 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
3.6%
Claims to check · rank 16 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
90.9%
Conflicts caught · rank 17 of 18
Unit
% of conflicts, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
94.1%
Content quality · rank 17 of 18
Unit
% of checks, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0080
Cost per product · rank 6 of 18
Unit
US dollars, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
92.6%
Decision accuracy · rank 16 of 18
Unit
% of decisions, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
91.7%
Field accuracy · rank 15 of 18
Unit
% of missing fields, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Invented values · rank 1 of 18
Unit
% of filled values, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Missing UK information · rank 1 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
14.3%
Not findable · rank 17 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
29.8%
Publish-ready listings · rank 16 of 18
Unit
% of products, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
17.9%
Reliably publish-ready · rank 15 of 18
Unit
% of products, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
49.4%
Unsupported claims · rank 16 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.74
Unsupported claims · rank 16 of 18
Unit
claims per product, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
30.9%
Wrong attributes · rank 16 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
11.3%
Wrong category or variant · rank 18 of 18
Unit
% of products, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · ROASBench
28.5
Audience score · rank 17 of 17
Unit
score out of 100, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
3.4
Business score · rank 17 of 17
Unit
score out of 100, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
8.7
Consistency score · rank 16 of 17
Unit
score out of 100, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
−$159,882
Contribution profit · rank 16 of 17
Unit
simulated US dollars, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
$0.15
Cost of a run · rank 5 of 17
Unit
US dollars, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
3
Months over budget · rank 17 of 17
Unit
months of 12, lower is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
15.2
Overall score · rank 17 of 17
Unit
score out of 100, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
55.2
Planning score · rank 1 of 17
Unit
score out of 100, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
−$0.17
Return on ad spend · rank 16 of 17
Unit
profit per $1 spent, higher is better
Configuration
Mistral Medium 3.5, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

-0.16
Confirmed task success · rank 40 of 46
Unit
IPS effect estimate, higher is better
Range
-0.20 to -0.13
Sample
6429 observations
Configuration
Mistral Medium 3.5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.15
Praise over complaint · rank 38 of 46
Unit
IPS effect estimate, higher is better
Range
-0.19 to -0.11
Sample
2310 observations
Configuration
Mistral Medium 3.5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.07
Steerability · rank 31 of 46
Unit
IPS effect estimate, higher is better
Range
-0.10 to -0.04
Sample
7189 observations
Configuration
Mistral Medium 3.5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.02
Tool grounding · rank 44 of 46
Unit
IPS effect estimate, higher is better
Range
-0.03 to -0.02
Sample
357598 observations
Configuration
Mistral Medium 3.5
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,430
Overall · rank 84 of 177
Unit
Arena rating, higher is better
Range
1,424 to 1,436
Sample
11632 votes
Configuration
Mistral Medium 3.5
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,430
Business, management and finance · rank 67 of 402
Unit
Arena rating, higher is better
Range
1,417 to 1,443
Sample
2256 votes
Configuration
Mistral Medium 3.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,398
Creative writing · rank 79 of 407
Unit
Arena rating, higher is better
Range
1,384 to 1,412
Sample
1999 votes
Configuration
Mistral Medium 3.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,434
Expert prompts · rank 93 of 359
Unit
Arena rating, higher is better
Range
1,416 to 1,451
Sample
1205 votes
Configuration
Mistral Medium 3.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,421
Instruction following · rank 77 of 409
Unit
Arena rating, higher is better
Range
1,411 to 1,431
Sample
3827 votes
Configuration
Mistral Medium 3.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,426
Overall · rank 95 of 409
Unit
Arena rating, higher is better
Range
1,420 to 1,433
Sample
11693 votes
Configuration
Mistral Medium 3.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,404
Writing, literature and language · rank 88 of 408
Unit
Arena rating, higher is better
Range
1,392 to 1,415
Sample
2919 votes
Configuration
Mistral Medium 3.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
46.9
Artificial Analysis Coding Index · rank 97 of 151
Unit
index score, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not your codebase or tools.
14.2
Artificial Analysis Intelligence Index · rank 224 of 314
Unit
index score, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not business work, and a blend: read the parts for any one task.
74.8%
GPQA Diamond · rank 190 of 276
Unit
% of questions, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not applied work; multiple-choice science questions.
13.8%
Humanity's Last Exam · rank 201 of 314
Unit
% of questions, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not everyday work; academic questions at the edge of expertise.
68.8%
IFBench · rank 58 of 215
Unit
% of instructions, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not judgement about what an instruction meant.
69.3%
Long-context reasoning (AA-LCR) · rank 175 of 309
Unit
% of questions, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not retrieval over your own document store.
40.2%
SciCode · rank 128 of 153
Unit
% of problems, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not general software engineering.
50.6%
Terminal-Bench 2.1 · rank 99 of 151
Unit
% of tasks, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not other harnesses or tools; superseded by 4.0 for newer models.
0.0%
Terminal-Bench 4.0 · rank 122 of 150
Unit
% of tasks, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not other harnesses or tools; one attempt per task.
33.3%
Terminal-Bench Hard · rank 80 of 213
Unit
% of tasks, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not other harnesses; Artificial Analysis no longer runs it on new models.
15.1%
Τ-bench banking · rank 98 of 142
Unit
% of tasks, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not your policies or systems; a simulated customer.
94.2%
Τ²-bench telecom · rank 28 of 214
Unit
% of tasks, higher is better
Configuration
Mistral Medium 3.5
Measured
1 Oct 2026
Not shown
Not your policies or systems; no longer run on new models.
Reported by UGI Leaderboard
60.0%
Requested-length error · rank 323 of 370
Unit
% off the requested word count, lower is better
Configuration
Mistral Medium 3.5 (high reasoning)
Measured
21 May 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
18.0%
Requested-length error · rank 150 of 370
Unit
% off the requested word count, lower is better
Configuration
Mistral Medium 3.5 (no reasoning)
Measured
21 May 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.31
Style adherence · rank 313 of 370
Unit
score from 0 to 1, higher is better
Configuration
Mistral Medium 3.5 (high reasoning)
Measured
21 May 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.35
Style adherence · rank 190 of 370
Unit
score from 0 to 1, higher is better
Configuration
Mistral Medium 3.5 (no reasoning)
Measured
21 May 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
43.9
Writing score · rank 185 of 370
Unit
score out of 100, higher is better
Configuration
Mistral Medium 3.5 (high reasoning)
Measured
21 May 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
45.5
Writing score · rank 174 of 370
Unit
score out of 100, higher is better
Configuration
Mistral Medium 3.5 (no reasoning)
Measured
21 May 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Mistral Medium 3.5 with