Compare / GPT-5.4 mini vs GPT-6.1 Sol
GPT-5.4 mini vs GPT-6.1 Sol
Which is cheaper, which is more capable, and which does better at each kind of work, from 23 results on 4 sources that measured both models.
Of the 23 results both models have, the better value on each. Results whose reported ranges overlap are too close to call.
Results as of 9 October 2026 (catalogue release 2026-10-09-dd620e778ac5).
GPT-5.4 mini or GPT-6.1 Sol? The short answer
Pick GPT-6.1 Sol: it scores higher (Intelligence Index 51.8 against 24.1) at about the same price.
- They cost about the same to run: $0.021 and $0.021 per game of bullet chess on BulletBench.
- GPT-6.1 Sol scores higher on the Artificial Analysis Intelligence Index: 51.8 against 24.1.
- GPT-6.1 Sol is newer: released 29 Sep 2026, against 17 Mar 2026 for GPT-5.4 mini.
- On our own benchmarks, on the results both models have: GPT-6.1 Sol does better for speed under pressure.
- On third-party benchmarks, on the results both models have: GPT-6.1 Sol does better for agents and tool use, reasoning and knowledge, writing and sticking to the facts.
Head to head
Every fact side by side
| GPT-5.4 mini | GPT-6.1 Sol | |
|---|---|---|
| Price per million tokens | $0.75 in · $4.50 out | $2.00 in · $10.00 out |
| Blended (3 in : 1 out) | $1.69 | $4.00 |
| Cheapest host, blended | $1.69 (Azure) | $4.00 (Azure) |
| Cost per game of bullet chess (BulletBench) | $0.021 | $0.021 |
| Cached input, per million | $0.075 cached input | $0.10 cached input |
| Intelligence Index | 24.1 | 51.8 |
| Spring Prompt overall | – | 81 |
| Output speed, tokens a second | – | 55 |
| Context window | 400,000 tokens | 1,050,000 tokens |
| Released | 17 Mar 2026 | 29 Sep 2026 (newer) |
| Weights | Closed | Closed |
| Developer based in | the US | the US |
Prices are the developer's list price where we have read it, otherwise the typical host on OpenRouter; see every model's price. Intelligence and speed from Artificial Analysis.
Which is better for what
Each model's rank among the models on sale today, for each kind of work, from our own benchmarks and licensed sources. Further right is better; the better rank is in green.
Every shared result
Each model's best result on every measure both have, grouped by benchmark, with the setting that produced it. Under each measure, every model that benchmark has measured is a grey tick, best to the right, so you can see where the two sit in the field. The headline measures of each benchmark come first; the rest are under "Show all". Green marks the better value; where a source reports ranges that overlap, neither is marked.
| Measure | GPT-5.4 mini | GPT-6.1 Sol |
|---|---|---|
| BulletBench measured by Spring Prompt · 3 results | ||
| Cost per game · Bullet 60sUS dollars, lower is better 21 · 22 of 28 measured | $0.0205no reasoning | $0.0210 |
| Ladder Elo · Bullet 60sladder Elo, higher is better 24 · 2 of 28 measured | 197no reasoning | 928Ultrafast (low reasoning) |
| Ladder Elo · Lightning 10+1ladder Elo, higher is better 7 · 5 of 23 measured | 571no reasoning | 584Ultrafast (low reasoning) |
| Arena (formerly LMArena) reported by Arena (formerly LMArena) · 3 results | ||
| Overall · TextArena rating, higher is better 72 · 18 of 387 measured | 1,447high reasoning | 1,483max reasoning |
| Overall · Text factualityArena rating, higher is better 64 · 37 of 163 measured | 1,447high reasoning | 1,463max reasoning |
| Business, management and finance · TextArena rating, higher is better 53 · 27 of 380 measured | 1,460high reasoning | 1,475max reasoning |
| Artificial Analysis reported by Artificial Analysis · 3 results | ||
| Artificial Analysis Intelligence Indexindex score, higher is better 81 · 6 of 243 measured | 24.1xhigh reasoning | 51.8max reasoning |
| Humanity's Last Exam% of questions, higher is better 82 · 8 of 242 measured | 28.1%xhigh reasoning | 52.9%max reasoning |
| Terminal-Bench 4.0% of tasks, higher is better 48 · 5 of 112 measured | 2.0%xhigh reasoning | 56.1%max reasoning |
| SimpleQA Verified (Epoch AI) reported by SimpleQA Verified (Epoch AI) · 1 result | ||
| Correct answers% of questions, higher is better 60 · 2 of 76 measured | 29.4%high reasoning | 73.9%max reasoning |
Show all 23 results (13 more)
| Measure | GPT-5.4 mini | GPT-6.1 Sol |
|---|---|---|
| BulletBench measured by Spring Prompt · 7 results | ||
| Cost per game · Lightning 10+1US dollars, lower is better 20 · 23 of 23 measured | $0.0136no reasoning | $0.16Ultrafast |
| Games lost on time · Bullet 60s% of games, lower is better 19 · 13 of 28 measured | 50.0%no reasoning | 25.0%Ultrafast (low reasoning) |
| Games lost on time · Lightning 10+1% of games, lower is better 1 · 16 of 23 measured | 0.0%no reasoning | 50.0%Ultrafast (low reasoning) |
| Invalid moves · Bullet 60s% of moves, lower is better 1 · 1 of 28 measured | 0.0%no reasoning | 0.0% |
| Invalid moves · Lightning 10+1% of moves, lower is better 1 · 1 of 23 measured | 0.0%no reasoning | 0.0%Ultrafast |
| Median move time · Bullet 60smilliseconds, lower is better 13 · 18 of 28 measured | 0.9 sno reasoning | 1.4 sUltrafast (low reasoning) |
| Median move time · Lightning 10+1milliseconds, lower is better 15 · 19 of 23 measured | 0.9 sno reasoning | 1.2 sUltrafast (low reasoning) |
| Arena (formerly LMArena) reported by Arena (formerly LMArena) · 4 results | ||
| Creative writing · TextArena rating, higher is better 92 · 21 of 385 measured | 1,401high reasoning | 1,460max reasoning |
| Expert prompts · TextArena rating, higher is better 63 · 4 of 338 measured | 1,480high reasoning | 1,543max reasoning |
| Instruction following · TextArena rating, higher is better 77 · 10 of 387 measured | 1,434high reasoning | 1,490max reasoning |
| Writing, literature and language · TextArena rating, higher is better 75 · 17 of 386 measured | 1,423high reasoning | 1,472max reasoning |
| Artificial Analysis reported by Artificial Analysis · 2 results | ||
| Long-context reasoning (AA-LCR)% of questions, higher is better 70 · 6 of 233 measured | 77.0%xhigh reasoning | 84.0%low reasoning |
| SciCode% of problems, higher is better 36 · 22 of 113 measured | 52.1%xhigh reasoning | 55.8%high reasoning |
Compare them with other models
- GPT-5.4 mini vs Muse Spark 1.3Most intelligent at $1–3 per million tokens
- GPT-5.4 mini vs Claude Opus 5.5Most intelligent available today
- GPT-6.1 Sol vs Claude Opus 5.5Most intelligent at $3–10 per million tokens
- GPT-6.1 Sol vs Muse Spark 1.3Cheapest with similar intelligence
- GPT-6.1 Sol vs GPT-6 AstraJust below overall
Quick answers
Is GPT-5.4 mini better than GPT-6.1 Sol?
Pick GPT-6.1 Sol: it scores higher (Intelligence Index 51.8 against 24.1) at about the same price.
Which is cheaper, GPT-5.4 mini or GPT-6.1 Sol?
They cost about the same to run: $0.021 and $0.021 per game of bullet chess on BulletBench.
Which is better for agents and tool use, GPT-5.4 mini or GPT-6.1 Sol?
GPT-6.1 Sol ranks 3 of 164 models on sale for multi-step tasks with tools: support desks, coding agents, function calls, against 56 for GPT-5.4 mini.
Which is better for reasoning and knowledge, GPT-5.4 mini or GPT-6.1 Sol?
GPT-6.1 Sol ranks 5 of 198 models on sale for hard questions across science, maths and general knowledge, against 71 for GPT-5.4 mini.
Which is better for writing, GPT-5.4 mini or GPT-6.1 Sol?
GPT-6.1 Sol ranks 15 of 188 models on sale for what people prefer in blind comparisons, and judged writing quality, against 58 for GPT-5.4 mini.
Which is better for sticking to the facts, GPT-5.4 mini or GPT-6.1 Sol?
GPT-6.1 Sol ranks 18 of 152 models on sale for summarising without inventing things, and factual answers, against 66 for GPT-5.4 mini.
Which is better for speed under pressure, GPT-5.4 mini or GPT-6.1 Sol?
GPT-6.1 Sol ranks 3 of 23 models on sale for good decisions against a real clock (fast chess), against 13 for GPT-5.4 mini.
Which has the bigger context window, GPT-5.4 mini or GPT-6.1 Sol?
GPT-6.1 Sol: 1,050,000 tokens, against 400,000 for GPT-5.4 mini.
Run GPT-5.4 mini and GPT-6.1 Sol on your own prompt
Benchmarks aren't your data. Try both side by side in the playground, with the cost of every answer.