Models / GPT-5.1-Codex

OpenAI

GPT-5.1-Codex

22nd most intelligent of the 34 models priced $3–10 per million tokens. Mid-field among models available today on every use case measured so far.

Price per million tokens
$1.25 in · $10.00 out
Typical host on OpenRouter, 7 Oct 2026 · compare prices
Speed
Not measured
Intelligence Index
23.7
Artificial Analysis
Spring Prompt overall
Not ranked yet
needs our benchmarks and two groups of results
Developer
OpenAI, based in the US
Weights not published
Released
Not recorded
400,000 tokens of context

Where it stands among the models you could choose

Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place. Retired models are left out.

AmongIntelligence IndexSpring Prompt overallPriceSpeed
Models available today 70th of 195Claude Opus 5.5 – 179th of 221Mistral Nemo –
Models priced $3–10 per million tokens 22nd of 34Claude Opus 5.5 – the group –
Models with similar intelligence the group – 33rd of 39MiMo-V2.5 –
Models from developers based in the US 42nd of 93Claude Opus 5.5 – 64th of 103Llama 3.1 8B Instruct –

The name under each rank is the leader of that group. Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.

How good is it, and for what?

Each use case ranks the available models its benchmarks measured, then those priced $3–10 per million tokens. The bar shows where it falls in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live Not measured
  • CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is Not measured
  • DeckBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending Not measured
  • ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls Not measured
  • tau2-bench: has not measured this model.
  • Berkeley Function Calling Leaderboard (BFCL) V4: has not measured this model.
  • OpenHands Index: has not measured this model.
  • Microsoft STATE-Bench: has not measured this model.
  • Artificial Analysis: has not measured this model.
  • Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects Not measured
  • APEX-Agents: has not measured this model.
  • GDP.pdf: has not measured this model.
  • Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge 75th of 195 available · 23rd of 34 at its price
BenchmarkGPT-5.1-CodexRank among availableBest available
Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported 23.7 70th of 195 Claude Opus 5.5 57.6
Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported 25.7% 78th of 194 Claude Opus 5.5 61.4%
Artificial AnalysisArtificial Analysis: GPQA Diamond · reported 86.0% 62nd of 185 GPT-6 Astra 96.3%
WritingWhat people prefer in blind comparisons, and judged writing quality Not measured
  • Arena (formerly LMArena): has not measured this model.
  • UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers Not measured
  • Vectara Hallucination Leaderboard: has not measured this model.
  • Arena (formerly LMArena): has not measured this model.
  • SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess) Not measured
  • BulletBench: has not measured this model.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

GPT-5.1-CodexClaude Opus 5.5
most intelligent at $3–10 per million tokens
MiMo-V2.5
cheapest with similar intelligence
GPT-6.1 Sol
leads overall
GPT-6 Astra
near the top overall
Price per million tokens $3.44 $8.00$0.21$4.00$20.00
Tokens a second – 97–5752
Intelligence Index 23.7 57.625.251.852.7
Spring Prompt overall – 82–8878
Rank among available models, by use case
Product listings – 6th–2nd1st
Decks from an analysis – 5th–3rd1st
Marketing planning – 5th––1st
Agents and tool use – 2nd95th3rd4th
Professional work – 3rd–9th5th
Reasoning and knowledge 75th of 195 1st74th5th3rd
Writing – 4th60th–17th
Sticking to the facts – 6th83rd2nd27th
Speed under pressure – –––14th

Shaded figures are better than GPT-5.1-Codex's. A dash means no result.

Where it has been measured

Every result

7 published results from 1 source, each in the source's own units, with the configuration that produced it.

Artificial Analysis · reported by Artificial Analysis · 7 results
MeasureValueRankConfigurationDated
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 23.7 159 of 402 GPT-5.1-Codex (high) 7 Oct 2026
Artificial Analysis: GPQA Diamond% of questions, higher is better 86.0% 113 of 359 GPT-5.1-Codex (high) 7 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 25.7% 159 of 400 GPT-5.1-Codex (high) 7 Oct 2026
Artificial Analysis: IFBench% of instructions, higher is better 70.0% 65 of 288 GPT-5.1-Codex (high) 7 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 69.3% 190 of 388 GPT-5.1-Codex (high) 7 Oct 2026
Artificial Analysis: Terminal-Bench Hard% of tasks, higher is better 34.8% 77 of 282 GPT-5.1-Codex (high) 7 Oct 2026
Artificial Analysis: τ²-bench telecom% of tasks, higher is better 83.0% 95 of 286 GPT-5.1-Codex (high) 7 Oct 2026

Not shown: Not business work, and a blend: read the parts for any one task.