Models / Claude Haiku 5.5

Anthropic

Claude Haiku 5.5

Claude Haiku 5.5 is best for reasoning and knowledge, and agents and tool use; a strong pick at a modest price.

Most intelligent of the 43 models priced under $0.25 per million tokens, and cheapest of the 17 models with similar intelligence. Among models available today, in the top quarter (by percentile) for reasoning and knowledge, and agents and tool use; in the bottom quarter for marketing planning. Results for professional work, writing and sticking to the facts are pending: those sources last updated before this model was released.

What a task costs, measured: ≈ $0.001 per product listing (CatalogBench) · ≈ $0.009 per deck (DeckBench) · ≈ $0.018 per survey plan and analysis (SurveyBench) · ≈ $0.016 per game of blitz chess (BulletBench)

Results as of 8 October 2026, from 119 results on 6 sources; prices checked 8 Oct 2026.

From our research: Claude Haiku 5.5 Is Far Better Than Haiku 4.5 at a Tenth of the Price. GPT-6 Luna Still Beats It at Business Work.

In the playground: our pick in 0 of 18 tests, 85% of checks passed →
Its best answers: Ball bouncing in a spinning hexagon (3 of 3 checks) · The solar system, to scale in time (3 of 3 checks) · A product page from one feed row (5 of 5 checks)

Price per million tokens
$0.10 in · $0.50 out
OpenRouter, 8 Oct 2026 · compare prices
Speed
235 tokens a second; on its test prompt the answer starts after 5.6 min, thinking time included
Artificial Analysis, on its usual API, at max reasoning; 143–235 across 5 settings
Intelligence Index
43.4
Artificial Analysis, at max reasoning; 29.4–43.4 across 5 settings
Spring Prompt overall
13th of 22 score 44 of 100 · how it works
Developer
Anthropic, based in the US
Weights not published
Released
7 Oct 2026
1,000,000 tokens of context

Where it stands among the models you could choose

Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place. Retired models are left out.

AmongIntelligence IndexSpring Prompt overallPrice, blendedSpeed
Models available today 17th of 198Claude Opus 5.5 13th of 22Claude Opus 5.5 $0.20 blended, 17th percentilecheapest: Mistral Nemo 5th of 64Trinity Large Thinking
Models priced under $0.25 per million tokens 1st of 43leads – sets this group 3rd of 18Nova Micro 1.0
Models with similar intelligence sets this group 5th of 7Gemini 3.8 Flash $0.20 blended, 0th percentilethe cheapest 1st of 11leads
Models writing 200 or more tokens a second 1st of 10leads 2nd of 3DeepSeek-V4.1-Flash $0.20 blended, 44th percentilecheapest: gpt-oss-20b sets this group
Models from developers based in the US 13th of 94Claude Opus 5.5 10th of 14Claude Opus 5.5 $0.20 blended, 17th percentilecheapest: Llama 3.1 8B Instruct 5th of 33Trinity Large Thinking

The name under each rank is the leader of that group; green is the top quarter of the group and red the bottom quarter, by rank. Price is shown as a percentile: the share of the group that costs less, so lower is cheaper. "Sets this group" marks the measure the group is defined by. Spring Prompt overall is our score out of 100 across every source (how it works). Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.

How good is it, and for what?

Each use case ranks the available models its benchmarks measured, then those priced under $0.25 per million tokens. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.

Product listingsTurning a sparse product feed and photos into listings that can go live 11th of 20 available
BenchmarkClaude Haiku 5.5Rank among availableBest available
CatalogBenchCatalogBench: reliably publish-ready 26.8% 14th of 20 GPT-6 Astra 75.0%
CatalogBenchCatalogBench (sales brief): reliably publish-ready 10.7% 9th of 20 GPT-6 Astra 69.6%
Decks from an analysisTurning a finished analysis into a deck you could present as it is 10th of 19 available
BenchmarkClaude Haiku 5.5Rank among availableBest available
DeckBenchDeckBench: deck rating 1,130 10th of 19 GPT-6 Astra 1,477
User surveysPlanning a user survey and reading its results without being misled 3rd of 8 available
BenchmarkClaude Haiku 5.5Rank among availableBest available
SurveyBenchSurveyBench: analysis rating 1,412 3rd of 8 GPT-6.1 Sol 1,650
Marketing planningPlanning a year of ad spend without overspending 17th of 19 available
BenchmarkClaude Haiku 5.5Rank among availableBest available
ROASBenchROASBench: overall score 24.1 17th of 19 GPT-6 Astra 56.6
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls 20th of 164 available · 2nd of 37 at its price
BenchmarkClaude Haiku 5.5Rank among availableBest available
Artificial AnalysisArtificial Analysis: Terminal-Bench 4.0 · reported 32.8% 15th of 105 Claude Sonnet 5.5 63.6%
  • tau2-bench: pending. Its latest results are dated 26 May 2026, before this model was released.
  • Berkeley Function Calling Leaderboard (BFCL) V4: pending. Its latest results are dated 16 Dec 2025, before this model was released.
  • OpenHands Index: pending. Its latest results are dated 30 Jun 2026, before this model was released.
  • Microsoft STATE-Bench: pending. Its latest results are dated 29 May 2026, before this model was released.
  • Vending-Bench 2: pending. Its latest results are dated 1 Oct 2026, before this model was released.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects PendingReleased after these sources last updated
  • APEX-Agents: pending. Its latest results are dated 1 Oct 2026, before this model was released.
  • GDP.pdf: pending. Its latest results are dated 1 Oct 2026, before this model was released.
  • Remote Labor Index: pending. Its latest results are dated 1 Oct 2026, before this model was released.
Reasoning and knowledgeHard questions across science, maths and general knowledge 19th of 198 available · 1st of 43 at its price
BenchmarkClaude Haiku 5.5Rank among availableBest available
Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported 43.4 17th of 198 Claude Opus 5.5 57.6
Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported 44.4% 20th of 197 Claude Opus 5.5 61.4%
WritingWhat people prefer in blind comparisons, and judged writing quality PendingReleased after these sources last updated
  • Arena (formerly LMArena): pending. Its latest results are dated 2 Oct 2026, before this model was released.
  • UGI Leaderboard: pending. Its latest results are dated 2 Oct 2026, before this model was released.
Sticking to the factsSummarising without inventing things, and factual answers PendingReleased after these sources last updated
  • Vectara Hallucination Leaderboard: pending. Its latest results are dated 22 Sep 2026, before this model was released.
  • Arena (formerly LMArena): pending. Its latest results are dated 2 Oct 2026, before this model was released.
  • SimpleQA Verified (Epoch AI): pending. Its latest results are dated 29 Sep 2026, before this model was released.
Speed under pressureGood decisions against a real clock (fast chess) 15th of 22 available · 2nd of 3 at its price
BenchmarkClaude Haiku 5.5Rank among availableBest available
BulletBenchBulletBench Bullet 60s: ladder Elo 310 14th of 20 Gemini 3.8 Flash 1,068

Nearest alternatives

  • Open weights at similar intelligenceGLM-5.3-Flash$0.24 blended · Intelligence Index 41.8

Against Claude Haiku 5.5 at $0.20 blended, Intelligence Index 43.4. "Similar intelligence" means within 4 points or better.

Against the alternatives

The models you would most likely weigh it against: the leaders of the groups above.

Scroll sideways to see every alternative.

Claude Haiku 5.5Claude Opus 5.5
most intelligent available today
Qwen3.8-Max (0902)
just above overall
Grok 4.7
just below overall
Price per million tokens $0.20 $8.00$3.00$3.00
Tokens a second 235 973764
Intelligence Index 43.4 57.645.446.4
Spring Prompt overall 44 854843
Rank among available models, by use case
Product listings 11th of 20 6th17th4th
Decks from an analysis 10th of 19 5th11th9th
User surveys 3rd of 8 2nd––
Marketing planning 17th of 19 5th14th11th
Agents and tool use 20th of 164 2nd7th27th
Professional work pending 3rd18th–
Reasoning and knowledge 19th of 198 1st17th16th
Writing pending 4th16th61st
Sticking to the facts pending 2nd38th106th
Speed under pressure 15th of 22 –––

Shaded figures are better than Claude Haiku 5.5's. A dash means no result.

Where it has been measured

Every result

119 published results from 6 sources, each in the source's own units, with the configuration that produced it.

BulletBench · measured by Spring Prompt · 10 results
MeasureValueRankConfigurationDated
BulletBench Blitz 3+2: cost per gameUS dollars, lower is better $0.0164 4 of 12 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
BulletBench Blitz 3+2: games lost on time% of games, lower is better 37.5% 6 of 12 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
BulletBench Blitz 3+2: invalid moves% of moves, lower is better 0.0% 1 of 12 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
BulletBench Blitz 3+2: ladder Eloladder Elo, higher is better 557
370–727
5 of 12 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
BulletBench Blitz 3+2: median move timemilliseconds, lower is better 6.2 s 8 of 12 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
BulletBench Bullet 60s: cost per gameUS dollars, lower is better $0.0019 6 of 31 Claude Haiku 5.5 (no reasoning)
none reasoning
8 Oct 2026
BulletBench Bullet 60s: games lost on time% of games, lower is better 37.5% 17 of 31 Claude Haiku 5.5 (no reasoning)
none reasoning
8 Oct 2026
BulletBench Bullet 60s: invalid moves% of moves, lower is better 0.0% 1 of 31 Claude Haiku 5.5 (no reasoning)
none reasoning
8 Oct 2026
BulletBench Bullet 60s: ladder Eloladder Elo, higher is better 310
0–575
23 of 31 Claude Haiku 5.5 (no reasoning)
none reasoning
8 Oct 2026
BulletBench Bullet 60s: median move timemilliseconds, lower is better 1.7 s 23 of 31 Claude Haiku 5.5 (no reasoning)
none reasoning
8 Oct 2026

Not shown: Not your cost: prices are those charged on the run date.

CatalogBench · measured by Spring Prompt · 30 results
MeasureValueRankConfigurationDated
CatalogBench (sales brief): channel rules broken% of products, lower is better 0.0% 1 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): claims to check% of products, lower is better 1.2% 12 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): content quality% of checks, higher is better 96.3% 13 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): failed outputs% of products, lower is better 0.0% 1 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): missing UK information% of products, lower is better 0.0% 1 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): not findable% of products, lower is better 10.7% 20 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): publish-ready listings% of products, higher is better 20.2% 11 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): reliably publish-ready% of products, higher is better 10.7% 9 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): unsupported claims% of products, lower is better 70.2% 12 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): unsupported claimsclaims per product, lower is better 2.46 12 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): wrong attributes% of products, lower is better 18.5% 10 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench (sales brief): wrong category or variant% of products, lower is better 0.0% 1 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: channel compliance% of products, higher is better 100.0% 1 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: channel rules broken% of products, lower is better 0.6% 11 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: claims to check% of products, lower is better 0.6% 13 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: conflicts caught% of conflicts, higher is better 100.0% 1 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: content quality% of checks, higher is better 95.6% 15 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: cost per productUS dollars, lower is better $0.0012 2 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: decision accuracy% of decisions, higher is better 95.6% 14 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: failed outputs% of products, lower is better 0.0% 1 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: field accuracy% of missing fields, higher is better 91.0% 18 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: invented values% of filled values, lower is better 0.8% 8 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: missing UK information% of products, lower is better 0.0% 1 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: not findable% of products, lower is better 11.9% 19 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: publish-ready listings% of products, higher is better 47.0% 13 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: reliably publish-ready% of products, higher is better 26.8% 14 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: unsupported claims% of products, lower is better 32.1% 13 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: unsupported claimsclaims per product, lower is better 1.20 16 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: wrong attributes% of products, lower is better 19.6% 13 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
CatalogBench: wrong category or variant% of products, lower is better 0.0% 1 of 21 Claude Haiku 5.5
provider default reasoning
8 Oct 2026

Not shown: Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.

DeckBench · measured by Spring Prompt · 16 results
MeasureValueRankConfigurationDated
DeckBench: accurate decks% of tasks, higher is better 33.3% 12 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: caveat dropped% of tasks, lower is better 16.7% 10 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: clean layout% of tasks, higher is better 66.7% 6 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: cost per deckUS dollars, lower is better $0.0091 2 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: deck ratingrating, higher is better 1,130
995–1,258
10 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: design quality% of the maximum, higher is better 61.1% 12 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: draft figure quoted% of tasks, lower is better 33.3% 11 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: findings missing% of tasks, lower is better 16.7% 14 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: head-to-head win rate% of comparisons, higher is better 50.0% 11 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: layout defects% of tasks, lower is better 33.3% 6 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: misleading metric used% of tasks, lower is better 16.7% 19 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: presentable decks% of tasks, higher is better 33.3% 5 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: recommendation late or wrong% of tasks, lower is better 33.3% 10 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: slides needing work% of tasks, lower is better 16.7% 2 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: unsupported claims% of tasks, lower is better 33.3% 13 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
DeckBench: unsupported numbers% of tasks, lower is better 16.7% 17 of 19 Claude Haiku 5.5
provider default reasoning
8 Oct 2026

Not shown: Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.

ROASBench · measured by Spring Prompt · 9 results
MeasureValueRankConfigurationDated
ROASBench: audience scorescore out of 100, higher is better 48.4 15 of 20 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
ROASBench: business scorescore out of 100, higher is better 13.5 16 of 20 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
ROASBench: consistency scorescore out of 100, higher is better 10.6 16 of 20 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
ROASBench: contribution profitsimulated US dollars, higher is better $32,146 16 of 20 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
ROASBench: cost of a runUS dollars, lower is better $0.0313 3 of 20 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
ROASBench: months over budgetmonths of 12, lower is better 0 1 of 20 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
ROASBench: overall scorescore out of 100, higher is better 24.1 18 of 20 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
ROASBench: planning scorescore out of 100, higher is better 54.8 11 of 20 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
ROASBench: return on ad spendprofit per $1 spent, higher is better $0.03 16 of 20 Claude Haiku 5.5
provider default reasoning
8 Oct 2026

Not shown: Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

SurveyBench · measured by Spring Prompt · 19 results
MeasureValueRankConfigurationDated
SurveyBench: analysis ratingrating, higher is better 1,412
1,322–1,492
3 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: analysis scorepoints of 100, higher is better 88.4 3 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: cost per taskUS dollars, lower is better $0.0185 2 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: head-to-head win rate% of comparisons, higher is better 72.5% 3 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: leading or double-barrelled questions% of tasks, lower is better 33.3% 3 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: mistaken next steps% of tasks, lower is better 22.2% 3 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: numbers right% of answers, higher is better 100.0% 1 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: objectives unanswerable% of tasks, lower is better 0.0% 1 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: precision ignored% of tasks, lower is better 0.0% 1 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: research scorepoints of 100, higher is better 90.2 3 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: sound research% of tasks, higher is better 11.1% 3 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: stakeholder's view presumed% of tasks, lower is better 0.0% 1 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: survey plan scorepoints of 100, higher is better 94.5 3 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: survey structure% of tasks, lower is better 22.2% 6 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: traps handled% of traps, higher is better 94.2% 3 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: traps missed% of tasks, lower is better 22.2% 1 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: unsupported findings% of tasks, lower is better 77.8% 4 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: wrong numbers% of tasks, lower is better 0.0% 1 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026
SurveyBench: wrong recommendation% of tasks, lower is better 11.1% 2 of 8 Claude Haiku 5.5
provider default reasoning
8 Oct 2026

Not shown: Not your customers or your data: invented organisations with simulated respondents, nine public tasks, one run per model per task at the provider's default reasoning setting.

Artificial Analysis · reported by Artificial Analysis · 35 results
MeasureValueRankConfigurationDated
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 37.8 75 of 407 Claude Haiku 5.5 (high reasoning) 8 Oct 2026
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 29.4 117 of 407 Claude Haiku 5.5 (low reasoning) 8 Oct 2026
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 43.4 42 of 407 Claude Haiku 5.5 (max reasoning) 8 Oct 2026
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 34.5 85 of 407 Claude Haiku 5.5 (medium reasoning) 8 Oct 2026
Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better 41.2 52 of 407 Claude Haiku 5.5 (xhigh reasoning) 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 37.3% 97 of 405 Claude Haiku 5.5 (high reasoning) 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 27.0% 160 of 405 Claude Haiku 5.5 (low reasoning) 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 44.4% 47 of 405 Claude Haiku 5.5 (max reasoning) 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 33.8% 119 of 405 Claude Haiku 5.5 (medium reasoning) 8 Oct 2026
Artificial Analysis: Humanity's Last Exam% of questions, higher is better 42.7% 57 of 405 Claude Haiku 5.5 (xhigh reasoning) 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 77.3% 117 of 393 Claude Haiku 5.5 (high reasoning) 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 72.0% 172 of 393 Claude Haiku 5.5 (low reasoning) 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 82.7% 30 of 393 Claude Haiku 5.5 (max reasoning) 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 77.3% 117 of 393 Claude Haiku 5.5 (medium reasoning) 8 Oct 2026
Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better 78.3% 100 of 393 Claude Haiku 5.5 (xhigh reasoning) 8 Oct 2026
Artificial Analysis: Output speedtokens per second, higher is better 143 34 of 133 Claude Haiku 5.5 (high reasoning) 8 Oct 2026
Artificial Analysis: Output speedtokens per second, higher is better 162 25 of 133 Claude Haiku 5.5 (low reasoning) 8 Oct 2026
Artificial Analysis: Output speedtokens per second, higher is better 235 5 of 133 Claude Haiku 5.5 (max reasoning) 8 Oct 2026
Artificial Analysis: Output speedtokens per second, higher is better 147 30 of 133 Claude Haiku 5.5 (medium reasoning) 8 Oct 2026
Artificial Analysis: Output speedtokens per second, higher is better 203 16 of 133 Claude Haiku 5.5 (xhigh reasoning) 8 Oct 2026
Artificial Analysis: SciCode% of problems, higher is better 48.7% 114 of 191 Claude Haiku 5.5 (high reasoning) 8 Oct 2026
Artificial Analysis: SciCode% of problems, higher is better 49.2% 110 of 191 Claude Haiku 5.5 (low reasoning) 8 Oct 2026
Artificial Analysis: SciCode% of problems, higher is better 55.0% 49 of 191 Claude Haiku 5.5 (max reasoning) 8 Oct 2026
Artificial Analysis: SciCode% of problems, higher is better 49.0% 113 of 191 Claude Haiku 5.5 (medium reasoning) 8 Oct 2026
Artificial Analysis: SciCode% of problems, higher is better 51.7% 83 of 191 Claude Haiku 5.5 (xhigh reasoning) 8 Oct 2026
Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better 21.7% 52 of 187 Claude Haiku 5.5 (high reasoning) 8 Oct 2026
Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better 12.6% 73 of 187 Claude Haiku 5.5 (low reasoning) 8 Oct 2026
Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better 32.8% 36 of 187 Claude Haiku 5.5 (max reasoning) 8 Oct 2026
Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better 15.2% 64 of 187 Claude Haiku 5.5 (medium reasoning) 8 Oct 2026
Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better 29.3% 42 of 187 Claude Haiku 5.5 (xhigh reasoning) 8 Oct 2026
Artificial Analysis: Time to first answer tokenseconds, lower is better 18.8 83 of 133 Claude Haiku 5.5 (high reasoning) 8 Oct 2026
Artificial Analysis: Time to first answer tokenseconds, lower is better 5.08 50 of 133 Claude Haiku 5.5 (low reasoning) 8 Oct 2026
Artificial Analysis: Time to first answer tokenseconds, lower is better 334 132 of 133 Claude Haiku 5.5 (max reasoning) 8 Oct 2026
Artificial Analysis: Time to first answer tokenseconds, lower is better 8.09 58 of 133 Claude Haiku 5.5 (medium reasoning) 8 Oct 2026
Artificial Analysis: Time to first answer tokenseconds, lower is better 63.5 120 of 133 Claude Haiku 5.5 (xhigh reasoning) 8 Oct 2026

Not shown: Not business work, and a blend: read the parts for any one task.