Models / Mistral Small 4

Mistral

Mistral Small 4

35 published results from 3 sources. Each card shows where the number comes from and what it does not measure. The overall leaderboard combines them; here each stands alone.

Provider
Mistral
Sources
3
Our benchmarks
1
Price per million tokens
$0.15 in · $0.60 out
OpenRouter list price, 1 Oct 2026 · 262,144-token context

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
$0.0039
Cost per game · rank 3 of 23
Unit
US dollars, lower is better
Configuration
Mistral Small 4 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
25.0%
Games lost on time · rank 7 of 23
Unit
% of games, lower is better
Configuration
Mistral Small 4 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.6%
Invalid moves · rank 20 of 23
Unit
% of moves, lower is better
Configuration
Mistral Small 4 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
367
Ladder Elo · rank 5 of 23
Unit
ladder Elo, higher is better
Range
126 to 564
Sample
12 games
Configuration
Mistral Small 4 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.6 s
Median move time · rank 5 of 23
Unit
milliseconds, lower is better
Configuration
Mistral Small 4 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · BulletBench
$0.0041
Cost per game · rank 4 of 14
Unit
US dollars, lower is better
Configuration
Mistral Small 4 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
0.0%
Games lost on time · rank 1 of 14
Unit
% of games, lower is better
Configuration
Mistral Small 4 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
1.0%
Invalid moves · rank 12 of 14
Unit
% of moves, lower is better
Configuration
Mistral Small 4 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
571
Ladder Elo · rank 1 of 14
Unit
ladder Elo, higher is better
Range
420 to 699
Sample
12 games
Configuration
Mistral Small 4 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.6 s
Median move time · rank 4 of 14
Unit
milliseconds, lower is better
Configuration
Mistral Small 4 (no reasoning), none reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.

Reported by others

26.6
Artificial Analysis Coding Index · rank 128 of 151
Unit
index score, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not your codebase or tools.
11.3
Artificial Analysis Intelligence Index · rank 254 of 314
Unit
index score, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not business work, and a blend: read the parts for any one task.
9.00
Artificial Analysis Intelligence Index · rank 280 of 314
Unit
index score, higher is better
Configuration
Mistral Small 4 (no reasoning)
Measured
1 Oct 2026
Not shown
Not business work, and a blend: read the parts for any one task.
76.9%
GPQA Diamond · rank 176 of 276
Unit
% of questions, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not applied work; multiple-choice science questions.
57.1%
GPQA Diamond · rank 253 of 276
Unit
% of questions, higher is better
Configuration
Mistral Small 4 (no reasoning)
Measured
1 Oct 2026
Not shown
Not applied work; multiple-choice science questions.
9.9%
Humanity's Last Exam · rank 231 of 314
Unit
% of questions, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not everyday work; academic questions at the edge of expertise.
3.8%
Humanity's Last Exam · rank 299 of 314
Unit
% of questions, higher is better
Configuration
Mistral Small 4 (no reasoning)
Measured
1 Oct 2026
Not shown
Not everyday work; academic questions at the edge of expertise.
48.2%
IFBench · rank 123 of 215
Unit
% of instructions, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not judgement about what an instruction meant.
32.8%
IFBench · rank 195 of 215
Unit
% of instructions, higher is better
Configuration
Mistral Small 4 (no reasoning)
Measured
1 Oct 2026
Not shown
Not judgement about what an instruction meant.
49.7%
Long-context reasoning (AA-LCR) · rank 231 of 309
Unit
% of questions, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not retrieval over your own document store.
28.3%
Long-context reasoning (AA-LCR) · rank 277 of 309
Unit
% of questions, higher is better
Configuration
Mistral Small 4 (no reasoning)
Measured
1 Oct 2026
Not shown
Not retrieval over your own document store.
38.8%
SciCode · rank 138 of 153
Unit
% of problems, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not general software engineering.
21.0%
Terminal-Bench 2.1 · rank 131 of 151
Unit
% of tasks, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not other harnesses or tools; superseded by 4.0 for newer models.
0.0%
Terminal-Bench 4.0 · rank 122 of 150
Unit
% of tasks, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not other harnesses or tools; one attempt per task.
17.4%
Terminal-Bench Hard · rank 143 of 213
Unit
% of tasks, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not other harnesses; Artificial Analysis no longer runs it on new models.
10.6%
Terminal-Bench Hard · rank 169 of 213
Unit
% of tasks, higher is better
Configuration
Mistral Small 4 (no reasoning)
Measured
1 Oct 2026
Not shown
Not other harnesses; Artificial Analysis no longer runs it on new models.
4.9%
Τ-bench banking · rank 136 of 142
Unit
% of tasks, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not your policies or systems; a simulated customer.
41.2%
Τ²-bench telecom · rank 151 of 214
Unit
% of tasks, higher is better
Configuration
Mistral Small 4
Measured
1 Oct 2026
Not shown
Not your policies or systems; no longer run on new models.
18.4%
Τ²-bench telecom · rank 207 of 214
Unit
% of tasks, higher is better
Configuration
Mistral Small 4 (no reasoning)
Measured
1 Oct 2026
Not shown
Not your policies or systems; no longer run on new models.
Reported by UGI Leaderboard
34.0%
Requested-length error · rank 269 of 370
Unit
% off the requested word count, lower is better
Configuration
Mistral Small 4 (high reasoning)
Measured
19 Mar 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
22.0%
Requested-length error · rank 202 of 370
Unit
% off the requested word count, lower is better
Configuration
Mistral Small 4 (no reasoning)
Measured
19 Mar 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.31
Style adherence · rank 310 of 370
Unit
score from 0 to 1, higher is better
Configuration
Mistral Small 4 (high reasoning)
Measured
19 Mar 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.29
Style adherence · rank 348 of 370
Unit
score from 0 to 1, higher is better
Configuration
Mistral Small 4 (no reasoning)
Measured
19 Mar 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
39.0
Writing score · rank 221 of 370
Unit
score out of 100, higher is better
Configuration
Mistral Small 4 (high reasoning)
Measured
19 Mar 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
40.3
Writing score · rank 213 of 370
Unit
score out of 100, higher is better
Configuration
Mistral Small 4 (no reasoning)
Measured
19 Mar 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Mistral Small 4 with