Models / Gemini 4 Argon

Google

Gemini 4 Argon

Gemini 4 Argon is best for writing and sticking to the facts; it is not on sale through OpenRouter, so it is here for reference.

Compared with the models available today (it is not on sale itself), in the top quarter (by percentile) for 4 of the 4 use cases measured, led by writing and sticking to the facts.

Results as of 8 October 2026, from 16 results on 2 sources; prices checked 8 Oct 2026.

Price per million tokens
Not available through OpenRouter
Speed
No speed yet
Artificial Analysis lists it but has not published a speed for it yet
Intelligence Index
52.6
Artificial Analysis, at high reasoning
Spring Prompt overall
Not ranked yet
needs our benchmarks and two groups of results
Developer
Google, based in the US
Weights not published
Released
30 Sep 2026 (Artificial Analysis)

Where it stands among the models you could choose

Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place. Retired models are left out.

AmongIntelligence IndexSpring Prompt overallPrice, blendedSpeed
Models available today 5th of 199Claude Opus 5.5 – – –
Models from developers based in the US 5th of 95Claude Opus 5.5 – – –

The name under each rank is the leader of that group; green is the top quarter of the group and red the bottom quarter, by rank. Price is shown as a percentile: the share of the group that costs less, so lower is cheaper. "Sets this group" marks the measure the group is defined by. Spring Prompt overall is our score out of 100 across every source (how it works). Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.

How good is it, and for what?

Each use case ranks the available models its benchmarks measured. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live Not measured
  • CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is Not measured
  • DeckBench: has not measured this model.
User surveysPlanning a user survey and reading its results without being misled Not measured
  • SurveyBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending Not measured
  • ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 3rd of 165 available
BenchmarkGemini 4 ArgonRank among availableBest available
Artificial AnalysisArtificial Analysis: Terminal-Bench 4.0 · reported 57.1% 4th of 106 Claude Sonnet 5.5 63.6%
  • tau2-bench: pending. Its latest results are dated 26 May 2026, before this model was released.
  • Berkeley Function Calling Leaderboard (BFCL) V4: pending. Its latest results are dated 16 Dec 2025, before this model was released.
  • OpenHands Index: pending. Its latest results are dated 30 Jun 2026, before this model was released.
  • Microsoft STATE-Bench: pending. Its latest results are dated 29 May 2026, before this model was released.
  • Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects Not measured
  • APEX-Agents: has not measured this model.
  • GDP.pdf: has not measured this model.
  • Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge 3rd of 199 available
BenchmarkGemini 4 ArgonRank among availableBest available
Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported 52.6 5th of 199 Claude Opus 5.5 57.6
Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported 57.1% 3rd of 198 Claude Opus 5.5 61.4%
WritingWhat people prefer in blind comparisons, and judged writing quality 1st of 189 available
BenchmarkGemini 4 ArgonRank among availableBest available
Arena (formerly LMArena)Arena Text: overall · reported 1,525 1st of 167 This model
  • UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers 1st of 153 available
BenchmarkGemini 4 ArgonRank among availableBest available
Arena (formerly LMArena)Arena Text factuality: overall · reported 1,511 1st of 114 This model
  • Vectara Hallucination Leaderboard: pending. Its latest results are dated 22 Sep 2026, before this model was released.
  • SimpleQA Verified (Epoch AI): pending. Its latest results are dated 29 Sep 2026, before this model was released.
Speed under pressureGood decisions against a real clock (fast chess) Not measured
  • BulletBench: has not measured this model.

Nearest alternatives

Against Gemini 4 Argon, Intelligence Index 52.6. "Similar intelligence" means within 4 points or better.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

Scroll sideways to see every alternative.

Gemini 4 ArgonClaude Opus 5.5
most intelligent available today
Gemini 3.8 Flash
Google's best other model
GPT-6.1 Sol
near the top overall
GPT-6 Astra
near the top overall
Price per million tokens – $8.00$3.00$4.00$20.00
Tokens a second – 971475552
Intelligence Index 52.6 57.640.951.852.7
Spring Prompt overall – 85618179
Rank among available models, by use case
Product listings – 6th12th2nd1st
Decks from an analysis – 5th12th3rd1st
User surveys – 2nd–1st–
Marketing planning – 5th7th–1st
Agents and tool use 3rd of 165 2nd32nd5th4th
Professional work – 3rd11th9th5th
Reasoning and knowledge 3rd of 199 1st13th6th4th
Writing 1st of 189 5th3rd16th20th
Sticking to the facts 1st of 153 3rd13th19th27th
Speed under pressure – –6th–14th

Shaded figures are better than Gemini 4 Argon's. A dash means no result.

Where it has been measured

Every result

16 published results from 2 sources, each in the source's own units, with the configuration that produced it.

Arena (formerly LMArena) · reported by Arena (formerly LMArena) · 11 results
MeasureValueRankConfigurationDated
Arena Agent: confirmed task successIPS effect estimate, higher is better 0.15
0.12–0.18
3 of 50 Gemini 4 Argon (high reasoning, arena agent) 2 Oct 2026
Arena Agent: praise over complaintIPS effect estimate, higher is better 0.28
0.20–0.35
4 of 50 Gemini 4 Argon (high reasoning, arena agent) 2 Oct 2026
Arena Agent: steerabilityIPS effect estimate, higher is better 0.13
0.12–0.15
1 of 50 Gemini 4 Argon (high reasoning, arena agent) 2 Oct 2026
Arena Agent: tool groundingIPS effect estimate, higher is better 0.00
-0.00–0.00
31 of 50 Gemini 4 Argon (high reasoning, arena agent) 2 Oct 2026
Arena Text factuality: overallArena rating, higher is better 1,511
1,503–1,518
1 of 182 Gemini 4 Argon (high reasoning) 2 Oct 2026
Arena Text: business, management and financeArena rating, higher is better 1,522
1,502–1,543
1 of 406 Gemini 4 Argon (high reasoning) 2 Oct 2026
Arena Text: creative writingArena rating, higher is better 1,519
1,499–1,538
1 of 411 Gemini 4 Argon (high reasoning) 2 Oct 2026
Arena Text: expert promptsArena rating, higher is better 1,539
1,513–1,565
7 of 363 Gemini 4 Argon (high reasoning) 2 Oct 2026
Arena Text: instruction followingArena rating, higher is better 1,528
1,514–1,543
1 of 413 Gemini 4 Argon (high reasoning) 2 Oct 2026
Arena Text: overallArena rating, higher is better 1,525
1,516–1,534
1 of 413 Gemini 4 Argon (high reasoning) 2 Oct 2026
Arena Text: writing, literature and languageArena rating, higher is better 1,522
1,505–1,539
1 of 412 Gemini 4 Argon (high reasoning) 2 Oct 2026

Not shown: Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.

Artificial Analysis · reported by Artificial Analysis · 5 results
MeasureValueRankConfigurationDated
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 52.6 8 of 407 Gemini 4 Argon (high reasoning) 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 57.1% 5 of 405 Gemini 4 Argon (high reasoning) 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 79.7% 78 of 393 Gemini 4 Argon (high reasoning) 8 Oct 2026
Artificial Analysis: SciCode% of problems, higher is better 61.8% 4 of 191 Gemini 4 Argon (high reasoning) 8 Oct 2026
Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better 57.1% 6 of 187 Gemini 4 Argon (high reasoning) 8 Oct 2026

Not shown: Not business work, and a blend: read the parts for any one task.