Compare / GPT-5.4 nano vs GPT-6.1 Sol
GPT-5.4 nano vs GPT-6.1 Sol
Which is cheaper, which is more capable, and which does better at each kind of work, from 23 results on 4 sources that measured both models.
Of the 23 results both models have, the better value on each. Results whose reported ranges overlap are too close to call.
Results as of 9 October 2026 (catalogue release 2026-10-09-dd620e778ac5).
GPT-5.4 nano or GPT-6.1 Sol? The short answer
Pick GPT-6.1 Sol if quality matters most: it scores higher (Intelligence Index 51.8 against 20.7). Pick GPT-5.4 nano if cost matters more: it costs 4.2 times less per game of bullet chess on BulletBench.
- GPT-5.4 nano is cheaper to run: $0.005 against $0.021 per game of bullet chess on BulletBench (4.2 times less).
- GPT-6.1 Sol scores higher on the Artificial Analysis Intelligence Index: 51.8 against 20.7.
- GPT-6.1 Sol is newer: released 29 Sep 2026, against 17 Mar 2026 for GPT-5.4 nano.
- On our own benchmarks, on the results both models have: GPT-6.1 Sol does better for speed under pressure.
- On third-party benchmarks, on the results both models have: GPT-6.1 Sol does better for agents and tool use, reasoning and knowledge, writing and sticking to the facts.
Head to head
Every fact side by side
| GPT-5.4 nano | GPT-6.1 Sol | |
|---|---|---|
| Price per million tokens | $0.20 in · $1.25 out | $2.00 in · $10.00 out |
| Blended (3 in : 1 out) | $0.46 | $4.00 |
| Cheapest host, blended | $0.46 (Azure) | $4.00 (Azure) |
| Cost per game of bullet chess (BulletBench) | $0.005 | $0.021 |
| Cached input, per million | $0.020 cached input | $0.10 cached input |
| Intelligence Index | 20.7 | 51.8 |
| Spring Prompt overall | – | 81 |
| Output speed, tokens a second | – | 55 |
| Context window | 400,000 tokens | 1,050,000 tokens |
| Released | 17 Mar 2026 | 29 Sep 2026 (newer) |
| Weights | Closed | Closed |
| Developer based in | the US | the US |
Prices are the developer's list price where we have read it, otherwise the typical host on OpenRouter; see every model's price. Intelligence and speed from Artificial Analysis.
Which is better for what
Each model's rank among the models on sale today, for each kind of work, from our own benchmarks and licensed sources. Further right is better; the better rank is in green.
Every shared result
Each model's best result on every measure both have, grouped by benchmark, with the setting that produced it. Under each measure, every model that benchmark has measured is a grey tick, best to the right, so you can see where the two sit in the field. The headline measures of each benchmark come first; the rest are under "Show all". Green marks the better value; where a source reports ranges that overlap, neither is marked.
| Measure | GPT-5.4 nano | GPT-6.1 Sol |
|---|---|---|
| BulletBench measured by Spring Prompt · 3 results | ||
| Cost per game · Bullet 60sUS dollars, lower is better 11 · 22 of 28 measured | $0.0050no reasoning | $0.0210 |
| Ladder Elo · Bullet 60sladder Elo, higher is better 25 · 2 of 28 measured | 33no reasoning | 928Ultrafast (low reasoning) |
| Ladder Elo · Lightning 10+1ladder Elo, higher is better 18 · 5 of 23 measured | 197no reasoning | 584Ultrafast (low reasoning) |
| Arena (formerly LMArena) reported by Arena (formerly LMArena) · 3 results | ||
| Overall · TextArena rating, higher is better 137 · 18 of 387 measured | 1,401high reasoning | 1,483max reasoning |
| Overall · Text factualityArena rating, higher is better 117 · 37 of 163 measured | 1,418high reasoning | 1,463max reasoning |
| Business, management and finance · TextArena rating, higher is better 132 · 27 of 380 measured | 1,405high reasoning | 1,475max reasoning |
| Artificial Analysis reported by Artificial Analysis · 3 results | ||
| Artificial Analysis Intelligence Indexindex score, higher is better 101 · 6 of 243 measured | 20.7xhigh reasoning | 51.8max reasoning |
| Humanity's Last Exam% of questions, higher is better 81 · 8 of 242 measured | 28.3%xhigh reasoning | 52.9%max reasoning |
| Terminal-Bench 4.0% of tasks, higher is better 58 · 5 of 112 measured | 0.5%xhigh reasoning | 56.1%max reasoning |
| SimpleQA Verified (Epoch AI) reported by SimpleQA Verified (Epoch AI) · 1 result | ||
| Correct answers% of questions, higher is better 72 · 2 of 76 measured | 11.7%high reasoning | 73.9%max reasoning |
Show all 23 results (13 more)
| Measure | GPT-5.4 nano | GPT-6.1 Sol |
|---|---|---|
| BulletBench measured by Spring Prompt · 7 results | ||
| Cost per game · Lightning 10+1US dollars, lower is better 8 · 23 of 23 measured | $0.0025no reasoning | $0.16Ultrafast |
| Games lost on time · Bullet 60s% of games, lower is better 24 · 13 of 28 measured | 66.7%no reasoning | 25.0%Ultrafast (low reasoning) |
| Games lost on time · Lightning 10+1% of games, lower is better 19 · 16 of 23 measured | 75.0%no reasoning | 50.0%Ultrafast (low reasoning) |
| Invalid moves · Bullet 60s% of moves, lower is better 27 · 1 of 28 measured | 3.1%no reasoning | 0.0% |
| Invalid moves · Lightning 10+1% of moves, lower is better 22 · 1 of 23 measured | 4.0%no reasoning | 0.0%Ultrafast |
| Median move time · Bullet 60smilliseconds, lower is better 15 · 18 of 28 measured | 1.1 sno reasoning | 1.4 sUltrafast (low reasoning) |
| Median move time · Lightning 10+1milliseconds, lower is better 17 · 19 of 23 measured | 1.0 sno reasoning | 1.2 sUltrafast (low reasoning) |
| Arena (formerly LMArena) reported by Arena (formerly LMArena) · 4 results | ||
| Creative writing · TextArena rating, higher is better 170 · 21 of 385 measured | 1,336high reasoning | 1,460max reasoning |
| Expert prompts · TextArena rating, higher is better 117 · 4 of 338 measured | 1,437high reasoning | 1,543max reasoning |
| Instruction following · TextArena rating, higher is better 139 · 10 of 387 measured | 1,387high reasoning | 1,490max reasoning |
| Writing, literature and language · TextArena rating, higher is better 156 · 17 of 386 measured | 1,363high reasoning | 1,472max reasoning |
| Artificial Analysis reported by Artificial Analysis · 2 results | ||
| Long-context reasoning (AA-LCR)% of questions, higher is better 71 · 6 of 233 measured | 76.7%xhigh reasoning | 84.0%low reasoning |
| SciCode% of problems, higher is better 56 · 22 of 113 measured | 47.2%xhigh reasoning | 55.8%high reasoning |
Compare them with other models
Quick answers
Is GPT-5.4 nano better than GPT-6.1 Sol?
Pick GPT-6.1 Sol if quality matters most: it scores higher (Intelligence Index 51.8 against 20.7). Pick GPT-5.4 nano if cost matters more: it costs 4.2 times less per game of bullet chess on BulletBench.
Which is cheaper, GPT-5.4 nano or GPT-6.1 Sol?
GPT-5.4 nano is cheaper to run: $0.005 against $0.021 per game of bullet chess on BulletBench (4.2 times less).
Which is better for agents and tool use, GPT-5.4 nano or GPT-6.1 Sol?
GPT-6.1 Sol ranks 3 of 164 models on sale for multi-step tasks with tools: support desks, coding agents, function calls, against 64 for GPT-5.4 nano.
Which is better for reasoning and knowledge, GPT-5.4 nano or GPT-6.1 Sol?
GPT-6.1 Sol ranks 5 of 198 models on sale for hard questions across science, maths and general knowledge, against 89 for GPT-5.4 nano.
Which is better for writing, GPT-5.4 nano or GPT-6.1 Sol?
GPT-6.1 Sol ranks 15 of 188 models on sale for what people prefer in blind comparisons, and judged writing quality, against 110 for GPT-5.4 nano.
Which is better for sticking to the facts, GPT-5.4 nano or GPT-6.1 Sol?
GPT-6.1 Sol ranks 18 of 152 models on sale for summarising without inventing things, and factual answers, against 89 for GPT-5.4 nano.
Which is better for speed under pressure, GPT-5.4 nano or GPT-6.1 Sol?
GPT-6.1 Sol ranks 3 of 23 models on sale for good decisions against a real clock (fast chess), against 19 for GPT-5.4 nano.
Which has the bigger context window, GPT-5.4 nano or GPT-6.1 Sol?
GPT-6.1 Sol: 1,050,000 tokens, against 400,000 for GPT-5.4 nano.
Run GPT-5.4 nano and GPT-6.1 Sol on your own prompt
Benchmarks aren't your data. Try both side by side in the playground, with the cost of every answer.