Compare / Mistral Small 4 vs GPT-6.1 Sol

Mistral Small 4 vs GPT-6.1 Sol

Which is cheaper, which is more capable, and which does better at each kind of work, from 20 results on 3 sources that measured both models.

Mistral AIMistral Small 4 Released 16 Mar 2026 · $0.15 in · $0.60 out
OpenAIGPT-6.1 Sol Released 29 Sep 2026 · $2.00 in · $10.00 out
Mistral Small 4 wins 72 too close to callGPT-6.1 Sol wins 11

Of the 20 results both models have, the better value on each. Results whose reported ranges overlap are too close to call.

Results as of 9 October 2026 (catalogue release 2026-10-09-dd620e778ac5). See also Mistral AI vs OpenAI, every model from both labs.

Mistral Small 4 or GPT-6.1 Sol? The short answer

Pick GPT-6.1 Sol if quality matters most: it scores higher (Intelligence Index 51.8 against 11.3). Pick Mistral Small 4 if cost matters more: it costs 5.4 times less per game of bullet chess on BulletBench. Mistral Small 4 is also faster.

  • Mistral Small 4 is cheaper to run: $0.004 against $0.021 per game of bullet chess on BulletBench (5.4 times less).
  • GPT-6.1 Sol scores higher on the Artificial Analysis Intelligence Index: 51.8 against 11.3.
  • Mistral Small 4 writes faster: 164 tokens a second against 55.
  • GPT-6.1 Sol is newer: released 29 Sep 2026, against 16 Mar 2026 for Mistral Small 4.
  • On our own benchmarks, on the results both models have: GPT-6.1 Sol does better for speed under pressure.
  • On third-party benchmarks, on the results both models have: GPT-6.1 Sol does better for agents and tool use, reasoning and knowledge and writing.

Head to head

11.3
Intelligence Index
51.8
$0.26
Price per million tokens, blendedlower is better
$4.00
164
Output speed, tokens a second
55
262,144
Context window, tokens
1,050,000

Every fact side by side

Mistral Small 4GPT-6.1 Sol
Price per million tokens$0.15 in · $0.60 out$2.00 in · $10.00 out
Blended (3 in : 1 out)$0.26$4.00
Cheapest host, blended$0.26 (Mistral)$4.00 (Azure)
Cost per game of bullet chess (BulletBench)$0.004$0.021
Cached input, per million$0.015 cached input$0.10 cached input
Intelligence Index11.351.8
Spring Prompt overall–81
Output speed, tokens a second16455
Context window262,144 tokens1,050,000 tokens
Released16 Mar 202629 Sep 2026 (newer)
WeightsOpenClosed
Developer based inFrancethe US

Prices are the developer's list price where we have read it, otherwise the typical host on OpenRouter; see every model's price. Intelligence and speed from Artificial Analysis.

Which is better for what

Each model's rank among the models on sale today, for each kind of work, from our own benchmarks and licensed sources. Further right is better; the better rank is in green.

Product listingsTurning a sparse product feed and photos into listings that can go live not measured 2 of 20
Decks from an analysisTurning a finished analysis into a deck you could present as it is not measured 3 of 19
User surveysPlanning a user survey and reading its results without being misled not measured 1 of 8
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 132 of 164 3 of 164
Professional workReal tasks from banking, consulting and law, business documents and freelance projects not measured 9 of 51
Reasoning and knowledgeHard questions across science, maths and general knowledge 131 of 198 5 of 198
WritingWhat people prefer in blind comparisons, and judged writing quality 116 of 188 15 of 188
Sticking to the factsSummarising without inventing things, and factual answers not measured 18 of 152
Speed under pressureGood decisions against a real clock (fast chess) 10 of 23 3 of 23

Every shared result

Each model's best result on every measure both have, grouped by benchmark, with the setting that produced it. Under each measure, every model that benchmark has measured is a grey tick, best to the right, so you can see where the two sit in the field. The headline measures of each benchmark come first; the rest are under "Show all". Green marks the better value; where a source reports ranges that overlap, neither is marked.

MeasureMistral Small 4GPT-6.1 Sol
BulletBench measured by Spring Prompt · 3 results
Cost per game · Bullet 60sUS dollars, lower is better 10 · 22 of 28 measured $0.0039no reasoning $0.0210
Ladder Elo · Bullet 60sladder Elo, higher is better 20 · 2 of 28 measured 367no reasoning 928Ultrafast (low reasoning)
Ladder Elo · Lightning 10+1ladder Elo, higher is better 7 · 5 of 23 measured 571no reasoning 584Ultrafast (low reasoning)
Artificial Analysis reported by Artificial Analysis · 3 results
Artificial Analysis Intelligence Indexindex score, higher is better 159 · 6 of 243 measured 11.3 51.8max reasoning
Humanity's Last Exam% of questions, higher is better 156 · 8 of 242 measured 9.9% 52.9%max reasoning
Terminal-Bench 4.0% of tasks, higher is better 66 · 5 of 112 measured 0.0% 56.1%max reasoning
UGI Leaderboard reported by UGI Leaderboard · 3 results
Writing score · UGIscore out of 100, higher is better 110 · 17 of 230 measured 40.3no reasoning 67.7high reasoning
Requested-length error · UGI% off the requested word count, lower is better 129 · 1 of 230 measured 22.0%no reasoning 0.0%high reasoning
Style adherence · UGIscore from 0 to 1, higher is better 192 · 39 of 230 measured 0.31high reasoning 0.37high reasoning
Show all 20 results (11 more)
MeasureMistral Small 4GPT-6.1 Sol
BulletBench measured by Spring Prompt · 7 results
Cost per game · Lightning 10+1US dollars, lower is better 11 · 23 of 23 measured $0.0041no reasoning $0.16Ultrafast
Games lost on time · Bullet 60s% of games, lower is better 13 · 13 of 28 measured 25.0%no reasoning 25.0%Ultrafast (low reasoning)
Games lost on time · Lightning 10+1% of games, lower is better 1 · 16 of 23 measured 0.0%no reasoning 50.0%Ultrafast (low reasoning)
Invalid moves · Bullet 60s% of moves, lower is better 26 · 1 of 28 measured 0.6%no reasoning 0.0%
Invalid moves · Lightning 10+1% of moves, lower is better 20 · 1 of 23 measured 1.0%no reasoning 0.0%Ultrafast
Median move time · Bullet 60smilliseconds, lower is better 8 · 18 of 28 measured 0.6 sno reasoning 1.4 sUltrafast (low reasoning)
Median move time · Lightning 10+1milliseconds, lower is better 7 · 19 of 23 measured 0.6 sno reasoning 1.2 sUltrafast (low reasoning)
Artificial Analysis reported by Artificial Analysis · 4 results
Long-context reasoning (AA-LCR)% of questions, higher is better 157 · 6 of 233 measured 49.7% 84.0%low reasoning
Output speedtokens per second, higher is better 16 · 58 of 74 measured 164 55.2max reasoning
SciCode% of problems, higher is better 83 · 22 of 113 measured 38.8% 55.8%high reasoning
Time to first answer tokenseconds, lower is better 6 · 36 of 74 measured 0.47no reasoning 1.69low reasoning

Quick answers

Is Mistral Small 4 better than GPT-6.1 Sol?

Pick GPT-6.1 Sol if quality matters most: it scores higher (Intelligence Index 51.8 against 11.3). Pick Mistral Small 4 if cost matters more: it costs 5.4 times less per game of bullet chess on BulletBench. Mistral Small 4 is also faster.

Which is cheaper, Mistral Small 4 or GPT-6.1 Sol?

Mistral Small 4 is cheaper to run: $0.004 against $0.021 per game of bullet chess on BulletBench (5.4 times less).

Which is faster, Mistral Small 4 or GPT-6.1 Sol?

Mistral Small 4 writes about 164 tokens a second on its usual API, against 55 for GPT-6.1 Sol (Artificial Analysis).

Which is better for agents and tool use, Mistral Small 4 or GPT-6.1 Sol?

GPT-6.1 Sol ranks 3 of 164 models on sale for multi-step tasks with tools: support desks, coding agents, function calls, against 132 for Mistral Small 4.

Which is better for reasoning and knowledge, Mistral Small 4 or GPT-6.1 Sol?

GPT-6.1 Sol ranks 5 of 198 models on sale for hard questions across science, maths and general knowledge, against 131 for Mistral Small 4.

Which is better for writing, Mistral Small 4 or GPT-6.1 Sol?

GPT-6.1 Sol ranks 15 of 188 models on sale for what people prefer in blind comparisons, and judged writing quality, against 116 for Mistral Small 4.

Which is better for speed under pressure, Mistral Small 4 or GPT-6.1 Sol?

GPT-6.1 Sol ranks 3 of 23 models on sale for good decisions against a real clock (fast chess), against 10 for Mistral Small 4.

Which has the bigger context window, Mistral Small 4 or GPT-6.1 Sol?

GPT-6.1 Sol: 1,050,000 tokens, against 262,144 for Mistral Small 4.

Run Mistral Small 4 and GPT-6.1 Sol on your own prompt

Benchmarks aren't your data. Try both side by side in the playground, with the cost of every answer.