Compare / GPT-5.4 mini vs GPT-6.1 Sol

GPT-5.4 mini vs GPT-6.1 Sol

Which is cheaper, which is more capable, and which does better at each kind of work, from 23 results on 4 sources that measured both models.

OpenAIGPT-5.4 mini Released 17 Mar 2026 · $0.75 in · $4.50 out
OpenAIGPT-6.1 Sol Released 29 Sep 2026 · $2.00 in · $10.00 out
GPT-5.4 mini wins 54 too close to callGPT-6.1 Sol wins 14

Of the 23 results both models have, the better value on each. Results whose reported ranges overlap are too close to call.

Results as of 9 October 2026 (catalogue release 2026-10-09-dd620e778ac5).

GPT-5.4 mini or GPT-6.1 Sol? The short answer

Pick GPT-6.1 Sol: it scores higher (Intelligence Index 51.8 against 24.1) at about the same price.

  • They cost about the same to run: $0.021 and $0.021 per game of bullet chess on BulletBench.
  • GPT-6.1 Sol scores higher on the Artificial Analysis Intelligence Index: 51.8 against 24.1.
  • GPT-6.1 Sol is newer: released 29 Sep 2026, against 17 Mar 2026 for GPT-5.4 mini.
  • On our own benchmarks, on the results both models have: GPT-6.1 Sol does better for speed under pressure.
  • On third-party benchmarks, on the results both models have: GPT-6.1 Sol does better for agents and tool use, reasoning and knowledge, writing and sticking to the facts.

Head to head

24.1
Intelligence Index
51.8
$1.69
Price per million tokens, blendedlower is better
$4.00
400,000
Context window, tokens
1,050,000

Every fact side by side

GPT-5.4 miniGPT-6.1 Sol
Price per million tokens$0.75 in · $4.50 out$2.00 in · $10.00 out
Blended (3 in : 1 out)$1.69$4.00
Cheapest host, blended$1.69 (Azure)$4.00 (Azure)
Cost per game of bullet chess (BulletBench)$0.021$0.021
Cached input, per million$0.075 cached input$0.10 cached input
Intelligence Index24.151.8
Spring Prompt overall–81
Output speed, tokens a second–55
Context window400,000 tokens1,050,000 tokens
Released17 Mar 202629 Sep 2026 (newer)
WeightsClosedClosed
Developer based inthe USthe US

Prices are the developer's list price where we have read it, otherwise the typical host on OpenRouter; see every model's price. Intelligence and speed from Artificial Analysis.

Which is better for what

Each model's rank among the models on sale today, for each kind of work, from our own benchmarks and licensed sources. Further right is better; the better rank is in green.

Product listingsTurning a sparse product feed and photos into listings that can go live not measured 2 of 20
Decks from an analysisTurning a finished analysis into a deck you could present as it is not measured 3 of 19
User surveysPlanning a user survey and reading its results without being misled not measured 1 of 8
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 56 of 164 3 of 164
Professional workReal tasks from banking, consulting and law, business documents and freelance projects not measured 9 of 51
Reasoning and knowledgeHard questions across science, maths and general knowledge 71 of 198 5 of 198
WritingWhat people prefer in blind comparisons, and judged writing quality 58 of 188 15 of 188
Sticking to the factsSummarising without inventing things, and factual answers 66 of 152 18 of 152
Speed under pressureGood decisions against a real clock (fast chess) 13 of 23 3 of 23

Every shared result

Each model's best result on every measure both have, grouped by benchmark, with the setting that produced it. Under each measure, every model that benchmark has measured is a grey tick, best to the right, so you can see where the two sit in the field. The headline measures of each benchmark come first; the rest are under "Show all". Green marks the better value; where a source reports ranges that overlap, neither is marked.

MeasureGPT-5.4 miniGPT-6.1 Sol
BulletBench measured by Spring Prompt · 3 results
Cost per game · Bullet 60sUS dollars, lower is better 21 · 22 of 28 measured $0.0205no reasoning $0.0210
Ladder Elo · Bullet 60sladder Elo, higher is better 24 · 2 of 28 measured 197no reasoning 928Ultrafast (low reasoning)
Ladder Elo · Lightning 10+1ladder Elo, higher is better 7 · 5 of 23 measured 571no reasoning 584Ultrafast (low reasoning)
Arena (formerly LMArena) reported by Arena (formerly LMArena) · 3 results
Overall · TextArena rating, higher is better 72 · 18 of 387 measured 1,447high reasoning 1,483max reasoning
Overall · Text factualityArena rating, higher is better 64 · 37 of 163 measured 1,447high reasoning 1,463max reasoning
Business, management and finance · TextArena rating, higher is better 53 · 27 of 380 measured 1,460high reasoning 1,475max reasoning
Artificial Analysis reported by Artificial Analysis · 3 results
Artificial Analysis Intelligence Indexindex score, higher is better 81 · 6 of 243 measured 24.1xhigh reasoning 51.8max reasoning
Humanity's Last Exam% of questions, higher is better 82 · 8 of 242 measured 28.1%xhigh reasoning 52.9%max reasoning
Terminal-Bench 4.0% of tasks, higher is better 48 · 5 of 112 measured 2.0%xhigh reasoning 56.1%max reasoning
SimpleQA Verified (Epoch AI) reported by SimpleQA Verified (Epoch AI) · 1 result
Correct answers% of questions, higher is better 60 · 2 of 76 measured 29.4%high reasoning 73.9%max reasoning
Show all 23 results (13 more)
MeasureGPT-5.4 miniGPT-6.1 Sol
BulletBench measured by Spring Prompt · 7 results
Cost per game · Lightning 10+1US dollars, lower is better 20 · 23 of 23 measured $0.0136no reasoning $0.16Ultrafast
Games lost on time · Bullet 60s% of games, lower is better 19 · 13 of 28 measured 50.0%no reasoning 25.0%Ultrafast (low reasoning)
Games lost on time · Lightning 10+1% of games, lower is better 1 · 16 of 23 measured 0.0%no reasoning 50.0%Ultrafast (low reasoning)
Invalid moves · Bullet 60s% of moves, lower is better 1 · 1 of 28 measured 0.0%no reasoning 0.0%
Invalid moves · Lightning 10+1% of moves, lower is better 1 · 1 of 23 measured 0.0%no reasoning 0.0%Ultrafast
Median move time · Bullet 60smilliseconds, lower is better 13 · 18 of 28 measured 0.9 sno reasoning 1.4 sUltrafast (low reasoning)
Median move time · Lightning 10+1milliseconds, lower is better 15 · 19 of 23 measured 0.9 sno reasoning 1.2 sUltrafast (low reasoning)
Arena (formerly LMArena) reported by Arena (formerly LMArena) · 4 results
Creative writing · TextArena rating, higher is better 92 · 21 of 385 measured 1,401high reasoning 1,460max reasoning
Expert prompts · TextArena rating, higher is better 63 · 4 of 338 measured 1,480high reasoning 1,543max reasoning
Instruction following · TextArena rating, higher is better 77 · 10 of 387 measured 1,434high reasoning 1,490max reasoning
Writing, literature and language · TextArena rating, higher is better 75 · 17 of 386 measured 1,423high reasoning 1,472max reasoning
Artificial Analysis reported by Artificial Analysis · 2 results
Long-context reasoning (AA-LCR)% of questions, higher is better 70 · 6 of 233 measured 77.0%xhigh reasoning 84.0%low reasoning
SciCode% of problems, higher is better 36 · 22 of 113 measured 52.1%xhigh reasoning 55.8%high reasoning

Quick answers

Is GPT-5.4 mini better than GPT-6.1 Sol?

Pick GPT-6.1 Sol: it scores higher (Intelligence Index 51.8 against 24.1) at about the same price.

Which is cheaper, GPT-5.4 mini or GPT-6.1 Sol?

They cost about the same to run: $0.021 and $0.021 per game of bullet chess on BulletBench.

Which is better for agents and tool use, GPT-5.4 mini or GPT-6.1 Sol?

GPT-6.1 Sol ranks 3 of 164 models on sale for multi-step tasks with tools: support desks, coding agents, function calls, against 56 for GPT-5.4 mini.

Which is better for reasoning and knowledge, GPT-5.4 mini or GPT-6.1 Sol?

GPT-6.1 Sol ranks 5 of 198 models on sale for hard questions across science, maths and general knowledge, against 71 for GPT-5.4 mini.

Which is better for writing, GPT-5.4 mini or GPT-6.1 Sol?

GPT-6.1 Sol ranks 15 of 188 models on sale for what people prefer in blind comparisons, and judged writing quality, against 58 for GPT-5.4 mini.

Which is better for sticking to the facts, GPT-5.4 mini or GPT-6.1 Sol?

GPT-6.1 Sol ranks 18 of 152 models on sale for summarising without inventing things, and factual answers, against 66 for GPT-5.4 mini.

Which is better for speed under pressure, GPT-5.4 mini or GPT-6.1 Sol?

GPT-6.1 Sol ranks 3 of 23 models on sale for good decisions against a real clock (fast chess), against 13 for GPT-5.4 mini.

Which has the bigger context window, GPT-5.4 mini or GPT-6.1 Sol?

GPT-6.1 Sol: 1,050,000 tokens, against 400,000 for GPT-5.4 mini.

Run GPT-5.4 mini and GPT-6.1 Sol on your own prompt

Benchmarks aren't your data. Try both side by side in the playground, with the cost of every answer.