Compare / GPT-6.1 Sol vs Grok 4.7
GPT-6.1 Sol vs Grok 4.7
Which is cheaper, which is more capable, and which does better at each kind of work, from 53 results on 3 sources that measured both models.
GPT-6.1 Sol or Grok 4.7? The short answer
- Grok 4.7 is cheaper: $3.00 against $4.00 per million tokens, 25% less.
- GPT-6.1 Sol scores higher on the Artificial Analysis Intelligence Index: 51.8 against 46.4.
- Grok 4.7 writes faster: 68 tokens a second against 57.
- On our business benchmarks, GPT-6.1 Sol ranks higher for product listings, decks from an analysis, agents and tool use, reasoning and knowledge and sticking to the facts.
At a glance
| GPT-6.1 Sol | Grok 4.7 | |
|---|---|---|
| Price per million tokens | $2.00 in · $10.00 out | $2.00 in · $6.00 out |
| Blended (3 in : 1 out) | $4.00 | $3.00 |
| Cheapest host, blended | $2.00 (OpenAI) | $3.00 (xAI) |
| Intelligence Index | 51.8 | 46.4 |
| Spring Prompt overall | 88 | 43 |
| Output speed, tokens a second | 57 | 68 |
| Context window | 1,050,000 tokens | 500,000 tokens |
| Weights | Closed | Closed |
| Developer based in | the US | the US |
Prices are the developer's list price where we have read it, otherwise the typical host on OpenRouter; see every model's price. Intelligence and speed from Artificial Analysis.
Which is better for what
Each model's rank among models available today, for each kind of work, from our own benchmarks and licensed sources. The better rank is in green.
| Use case | GPT-6.1 Sol | Grok 4.7 |
|---|---|---|
| Product listingsTurning a sparse product feed and photos into listings that can go live | 2 of 19 | 4 of 19 |
| Decks from an analysisTurning a finished analysis into a deck you could present as it is | 3 of 18 | 9 of 18 |
| Marketing planningPlanning a year of ad spend without overspending | not measured | 11 of 18 |
| Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls | 3 of 161 | 26 of 161 |
| Professional workReal tasks from banking, consulting and law, business documents and freelance projects | 9 of 51 | not measured |
| Reasoning and knowledgeHard questions across science, maths and general knowledge | 5 of 195 | 16 of 195 |
| WritingWhat people prefer in blind comparisons, and judged writing quality | not measured | 63 of 187 |
| Sticking to the factsSummarising without inventing things, and factual answers | 2 of 153 | 112 of 153 |
| Speed under pressureGood decisions against a real clock (fast chess) | not measured | 22 of 22 |
Their slides, side by side
The same brief for both: Manchester store investment review for Hearthside Coffee, an invented company. First slide of each deck from DeckBench.
Every shared result
| Benchmark | Metric | GPT-6.1 Sol | Grok 4.7 |
|---|---|---|---|
| CatalogBenchmeasured by us | Channel rules broken% of products | 0.0% | 0.0% |
| CatalogBenchmeasured by us | Claims to check% of products | 0.0% | 0.0% |
| CatalogBenchmeasured by us | Content quality% of checks | 98.8% | 98.1% |
| CatalogBenchmeasured by us | Failed outputs% of products | 0.0% | 0.0% |
| CatalogBenchmeasured by us | Missing UK information% of products | 0.0% | 0.0% |
| CatalogBenchmeasured by us | Not findable% of products | 5.4% | 4.2% |
| CatalogBenchmeasured by us | Publish-ready listings% of products | 72.0% | 58.9% |
| CatalogBenchmeasured by us | Reliably publish-ready% of products | 67.9% | 33.9% |
| CatalogBenchmeasured by us | Unsupported claimsclaims per product | 0.07 | 0.71 |
| CatalogBenchmeasured by us | Unsupported claims% of products | 3.6% | 26.8% |
| CatalogBenchmeasured by us | Wrong attributes% of products | 19.1% | 15.5% |
| CatalogBenchmeasured by us | Wrong category or variant% of products | 1.8% | 2.4% |
| CatalogBenchmeasured by us | Channel compliance% of products | 100.0% | 100.0% |
| CatalogBenchmeasured by us | Channel rules broken% of products | 0.0% | 0.0% |
| CatalogBenchmeasured by us | Claims to check% of products | 0.0% | 0.0% |
| CatalogBenchmeasured by us | Conflicts caught% of conflicts | 100.0% | 100.0% |
| CatalogBenchmeasured by us | Content quality% of checks | 98.6% | 97.6% |
| CatalogBenchmeasured by us | Cost per productUS dollars | $0.0084 | $0.0237 |
| CatalogBenchmeasured by us | Decision accuracy% of decisions | 98.7% | 97.3% |
| CatalogBenchmeasured by us | Failed outputs% of products | 0.0% | 0.0% |
| CatalogBenchmeasured by us | Field accuracy% of missing fields | 94.4% | 92.9% |
| CatalogBenchmeasured by us | Invented values% of filled values | 2.2% | 0.5% |
| CatalogBenchmeasured by us | Missing UK information% of products | 0.0% | 0.0% |
| CatalogBenchmeasured by us | Not findable% of products | 4.8% | 5.4% |
| CatalogBenchmeasured by us | Publish-ready listings% of products | 78.0% | 70.8% |
| CatalogBenchmeasured by us | Reliably publish-ready% of products | 75.0% | 57.1% |
| CatalogBenchmeasured by us | Unsupported claims% of products | 0.0% | 11.3% |
| CatalogBenchmeasured by us | Unsupported claimsclaims per product | 0.00 | 0.49 |
| CatalogBenchmeasured by us | Wrong attributes% of products | 16.7% | 15.5% |
| CatalogBenchmeasured by us | Wrong category or variant% of products | 1.8% | 2.4% |
| DeckBenchmeasured by us | Accurate decks% of tasks | 100.0% | 66.7% |
| DeckBenchmeasured by us | Caveat dropped% of tasks | 0.0% | 16.7% |
| DeckBenchmeasured by us | Clean layout% of tasks | 83.3% | 83.3% |
| DeckBenchmeasured by us | Cost per deckUS dollars | $0.15 | $0.25 |
| DeckBenchmeasured by us | Deck ratingrating | 1,412 | 1,242 |
| DeckBenchmeasured by us | Design quality% of the maximum | 71.4% | 76.4% |
| DeckBenchmeasured by us | Draft figure quoted% of tasks | 0.0% | 16.7% |
| DeckBenchmeasured by us | Findings missing% of tasks | 0.0% | 0.0% |
| DeckBenchmeasured by us | Head-to-head win rate% of comparisons | 74.6% | 68.7% |
| DeckBenchmeasured by us | Layout defects% of tasks | 16.7% | 16.7% |
| DeckBenchmeasured by us | Misleading metric used% of tasks | 0.0% | 0.0% |
| DeckBenchmeasured by us | Presentable decks% of tasks | 50.0% | 33.3% |
| DeckBenchmeasured by us | Recommendation late or wrong% of tasks | 0.0% | 16.7% |
| DeckBenchmeasured by us | Slides needing work% of tasks | 50.0% | 16.7% |
| DeckBenchmeasured by us | Unsupported claims% of tasks | 0.0% | 0.0% |
| DeckBenchmeasured by us | Unsupported numbers% of tasks | 0.0% | 0.0% |
| Artificial Analysisreported | Artificial Analysis Intelligence Indexindex score | 51.8 | 46.4 |
| Artificial Analysisreported | Humanity's Last Exam% of questions | 52.9% | 43.1% |
| Artificial Analysisreported | Long-context reasoning (AA-LCR)% of questions | 84.0% | 80.0% |
| Artificial Analysisreported | Output speedtokens per second | 57.2 | 68.4 |
| Artificial Analysisreported | SciCode% of problems | 55.8% | 57.8% |
| Artificial Analysisreported | Terminal-Bench 4.0% of tasks | 56.1% | 25.8% |
| Artificial Analysisreported | Time to first answer tokenseconds | 1.51 | 4.25 |
Bold green marks the better value on that metric. Where a source reports ranges that overlap, the difference may not be meaningful; see the benchmark page for ranges.
Quick answers
Is GPT-6.1 Sol better than Grok 4.7?
It depends on the task. Among models available today, on our benchmarks and licensed sources, GPT-6.1 Sol ranks higher for product listings, decks from an analysis, agents and tool use, reasoning and knowledge and sticking to the facts.
Which is cheaper, GPT-6.1 Sol or Grok 4.7?
Grok 4.7 costs $3.00 per million tokens (three input to one output), against $4.00 for GPT-6.1 Sol.
Which is faster, GPT-6.1 Sol or Grok 4.7?
Grok 4.7 writes about 68 tokens a second on its usual API, against 57 for GPT-6.1 Sol (Artificial Analysis).
Run GPT-6.1 Sol and Grok 4.7 on your own prompt
Benchmarks aren't your data. Try both side by side in the playground, with the cost of every answer.

