Compare / Mistral Small 4 vs GPT-6.1 Sol
Mistral Small 4 vs GPT-6.1 Sol
Which is cheaper, which is more capable, and which does better at each kind of work, from 20 results on 3 sources that measured both models.
Of the 20 results both models have, the better value on each. Results whose reported ranges overlap are too close to call.
Results as of 9 October 2026 (catalogue release 2026-10-09-dd620e778ac5). See also Mistral AI vs OpenAI, every model from both labs.
Mistral Small 4 or GPT-6.1 Sol? The short answer
Pick GPT-6.1 Sol if quality matters most: it scores higher (Intelligence Index 51.8 against 11.3). Pick Mistral Small 4 if cost matters more: it costs 5.4 times less per game of bullet chess on BulletBench. Mistral Small 4 is also faster.
- Mistral Small 4 is cheaper to run: $0.004 against $0.021 per game of bullet chess on BulletBench (5.4 times less).
- GPT-6.1 Sol scores higher on the Artificial Analysis Intelligence Index: 51.8 against 11.3.
- Mistral Small 4 writes faster: 164 tokens a second against 55.
- GPT-6.1 Sol is newer: released 29 Sep 2026, against 16 Mar 2026 for Mistral Small 4.
- On our own benchmarks, on the results both models have: GPT-6.1 Sol does better for speed under pressure.
- On third-party benchmarks, on the results both models have: GPT-6.1 Sol does better for agents and tool use, reasoning and knowledge and writing.
Head to head
Every fact side by side
| Mistral Small 4 | GPT-6.1 Sol | |
|---|---|---|
| Price per million tokens | $0.15 in · $0.60 out | $2.00 in · $10.00 out |
| Blended (3 in : 1 out) | $0.26 | $4.00 |
| Cheapest host, blended | $0.26 (Mistral) | $4.00 (Azure) |
| Cost per game of bullet chess (BulletBench) | $0.004 | $0.021 |
| Cached input, per million | $0.015 cached input | $0.10 cached input |
| Intelligence Index | 11.3 | 51.8 |
| Spring Prompt overall | – | 81 |
| Output speed, tokens a second | 164 | 55 |
| Context window | 262,144 tokens | 1,050,000 tokens |
| Released | 16 Mar 2026 | 29 Sep 2026 (newer) |
| Weights | Open | Closed |
| Developer based in | France | the US |
Prices are the developer's list price where we have read it, otherwise the typical host on OpenRouter; see every model's price. Intelligence and speed from Artificial Analysis.
Which is better for what
Each model's rank among the models on sale today, for each kind of work, from our own benchmarks and licensed sources. Further right is better; the better rank is in green.
Every shared result
Each model's best result on every measure both have, grouped by benchmark, with the setting that produced it. Under each measure, every model that benchmark has measured is a grey tick, best to the right, so you can see where the two sit in the field. The headline measures of each benchmark come first; the rest are under "Show all". Green marks the better value; where a source reports ranges that overlap, neither is marked.
| Measure | Mistral Small 4 | GPT-6.1 Sol |
|---|---|---|
| BulletBench measured by Spring Prompt · 3 results | ||
| Cost per game · Bullet 60sUS dollars, lower is better 10 · 22 of 28 measured | $0.0039no reasoning | $0.0210 |
| Ladder Elo · Bullet 60sladder Elo, higher is better 20 · 2 of 28 measured | 367no reasoning | 928Ultrafast (low reasoning) |
| Ladder Elo · Lightning 10+1ladder Elo, higher is better 7 · 5 of 23 measured | 571no reasoning | 584Ultrafast (low reasoning) |
| Artificial Analysis reported by Artificial Analysis · 3 results | ||
| Artificial Analysis Intelligence Indexindex score, higher is better 159 · 6 of 243 measured | 11.3 | 51.8max reasoning |
| Humanity's Last Exam% of questions, higher is better 156 · 8 of 242 measured | 9.9% | 52.9%max reasoning |
| Terminal-Bench 4.0% of tasks, higher is better 66 · 5 of 112 measured | 0.0% | 56.1%max reasoning |
| UGI Leaderboard reported by UGI Leaderboard · 3 results | ||
| Writing score · UGIscore out of 100, higher is better 110 · 17 of 230 measured | 40.3no reasoning | 67.7high reasoning |
| Requested-length error · UGI% off the requested word count, lower is better 129 · 1 of 230 measured | 22.0%no reasoning | 0.0%high reasoning |
| Style adherence · UGIscore from 0 to 1, higher is better 192 · 39 of 230 measured | 0.31high reasoning | 0.37high reasoning |
Show all 20 results (11 more)
| Measure | Mistral Small 4 | GPT-6.1 Sol |
|---|---|---|
| BulletBench measured by Spring Prompt · 7 results | ||
| Cost per game · Lightning 10+1US dollars, lower is better 11 · 23 of 23 measured | $0.0041no reasoning | $0.16Ultrafast |
| Games lost on time · Bullet 60s% of games, lower is better 13 · 13 of 28 measured | 25.0%no reasoning | 25.0%Ultrafast (low reasoning) |
| Games lost on time · Lightning 10+1% of games, lower is better 1 · 16 of 23 measured | 0.0%no reasoning | 50.0%Ultrafast (low reasoning) |
| Invalid moves · Bullet 60s% of moves, lower is better 26 · 1 of 28 measured | 0.6%no reasoning | 0.0% |
| Invalid moves · Lightning 10+1% of moves, lower is better 20 · 1 of 23 measured | 1.0%no reasoning | 0.0%Ultrafast |
| Median move time · Bullet 60smilliseconds, lower is better 8 · 18 of 28 measured | 0.6 sno reasoning | 1.4 sUltrafast (low reasoning) |
| Median move time · Lightning 10+1milliseconds, lower is better 7 · 19 of 23 measured | 0.6 sno reasoning | 1.2 sUltrafast (low reasoning) |
| Artificial Analysis reported by Artificial Analysis · 4 results | ||
| Long-context reasoning (AA-LCR)% of questions, higher is better 157 · 6 of 233 measured | 49.7% | 84.0%low reasoning |
| Output speedtokens per second, higher is better 16 · 58 of 74 measured | 164 | 55.2max reasoning |
| SciCode% of problems, higher is better 83 · 22 of 113 measured | 38.8% | 55.8%high reasoning |
| Time to first answer tokenseconds, lower is better 6 · 36 of 74 measured | 0.47no reasoning | 1.69low reasoning |
Compare them with other models
- Mistral AI vs OpenAIEvery model from both labs, best against best
- Mistral Small 4 vs Claude Opus 5.5Most intelligent available today
- GPT-6.1 Sol vs Claude Opus 5.5Most intelligent at $3–10 per million tokens
- GPT-6.1 Sol vs Muse Spark 1.3Cheapest with similar intelligence
- GPT-6.1 Sol vs GPT-6 AstraJust below overall
Quick answers
Is Mistral Small 4 better than GPT-6.1 Sol?
Pick GPT-6.1 Sol if quality matters most: it scores higher (Intelligence Index 51.8 against 11.3). Pick Mistral Small 4 if cost matters more: it costs 5.4 times less per game of bullet chess on BulletBench. Mistral Small 4 is also faster.
Which is cheaper, Mistral Small 4 or GPT-6.1 Sol?
Mistral Small 4 is cheaper to run: $0.004 against $0.021 per game of bullet chess on BulletBench (5.4 times less).
Which is faster, Mistral Small 4 or GPT-6.1 Sol?
Mistral Small 4 writes about 164 tokens a second on its usual API, against 55 for GPT-6.1 Sol (Artificial Analysis).
Which is better for agents and tool use, Mistral Small 4 or GPT-6.1 Sol?
GPT-6.1 Sol ranks 3 of 164 models on sale for multi-step tasks with tools: support desks, coding agents, function calls, against 132 for Mistral Small 4.
Which is better for reasoning and knowledge, Mistral Small 4 or GPT-6.1 Sol?
GPT-6.1 Sol ranks 5 of 198 models on sale for hard questions across science, maths and general knowledge, against 131 for Mistral Small 4.
Which is better for writing, Mistral Small 4 or GPT-6.1 Sol?
GPT-6.1 Sol ranks 15 of 188 models on sale for what people prefer in blind comparisons, and judged writing quality, against 116 for Mistral Small 4.
Which is better for speed under pressure, Mistral Small 4 or GPT-6.1 Sol?
GPT-6.1 Sol ranks 3 of 23 models on sale for good decisions against a real clock (fast chess), against 10 for Mistral Small 4.
Which has the bigger context window, Mistral Small 4 or GPT-6.1 Sol?
GPT-6.1 Sol: 1,050,000 tokens, against 262,144 for Mistral Small 4.
Run Mistral Small 4 and GPT-6.1 Sol on your own prompt
Benchmarks aren't your data. Try both side by side in the playground, with the cost of every answer.