Models / Claude Haiku 5.5
Claude Haiku 5.5
Claude Haiku 5.5 is best for reasoning and knowledge, and agents and tool use; a strong pick at a modest price.
Most intelligent of the 43 models priced under $0.25 per million tokens, and cheapest of the 17 models with similar intelligence. Among models available today, in the top quarter (by percentile) for reasoning and knowledge, and agents and tool use; in the bottom quarter for marketing planning. Results for professional work, writing and sticking to the facts are pending: those sources last updated before this model was released.
What a task costs, measured: ≈ $0.001 per product listing (CatalogBench) · ≈ $0.009 per deck (DeckBench) · ≈ $0.018 per survey plan and analysis (SurveyBench) · ≈ $0.016 per game of blitz chess (BulletBench)
Results as of 8 October 2026, from 119 results on 6 sources; prices checked 8 Oct 2026.
From our research: Claude Haiku 5.5 Is Far Better Than Haiku 4.5 at a Tenth of the Price. GPT-6 Luna Still Beats It at Business Work.
In the playground: our pick in 0 of 18 tests, 85% of checks passed →
Its best answers: Ball bouncing in a spinning hexagon (3 of 3 checks) · The solar system, to scale in time (3 of 3 checks) · A product page from one feed row (5 of 5 checks)
- Price per million tokens
- $0.10 in · $0.50 out
OpenRouter, 8 Oct 2026 · compare prices - Speed
- 235 tokens a second; on its test prompt the answer starts after 5.6 min, thinking time included
Artificial Analysis, on its usual API, at max reasoning; 143–235 across 5 settings - Intelligence Index
- 43.4
Artificial Analysis, at max reasoning; 29.4–43.4 across 5 settings - Spring Prompt overall
- 13th of 22 score 44 of 100 · how it works
- Developer
- Anthropic, based in the US
Weights not published - Released
- 7 Oct 2026
1,000,000 tokens of context
Where it stands among the models you could choose
Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place. Retired models are left out.
| Among | Intelligence Index | Spring Prompt overall | Price, blended | Speed |
|---|---|---|---|---|
| Models available today | 17th of 198Claude Opus 5.5 | 13th of 22Claude Opus 5.5 | $0.20 blended, 17th percentilecheapest: Mistral Nemo | 5th of 64Trinity Large Thinking |
| Models priced under $0.25 per million tokens | 1st of 43leads | – | sets this group | 3rd of 18Nova Micro 1.0 |
| Models with similar intelligence | sets this group | 5th of 7Gemini 3.8 Flash | $0.20 blended, 0th percentilethe cheapest | 1st of 11leads |
| Models writing 200 or more tokens a second | 1st of 10leads | 2nd of 3DeepSeek-V4.1-Flash | $0.20 blended, 44th percentilecheapest: gpt-oss-20b | sets this group |
| Models from developers based in the US | 13th of 94Claude Opus 5.5 | 10th of 14Claude Opus 5.5 | $0.20 blended, 17th percentilecheapest: Llama 3.1 8B Instruct | 5th of 33Trinity Large Thinking |
The name under each rank is the leader of that group; green is the top quarter of the group and red the bottom quarter, by rank. Price is shown as a percentile: the share of the group that costs less, so lower is cheaper. "Sets this group" marks the measure the group is defined by. Spring Prompt overall is our score out of 100 across every source (how it works). Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.
How good is it, and for what?
Each use case ranks the available models its benchmarks measured, then those priced under $0.25 per million tokens. The bar shows its percentile in that field, best to the right. Open a row for the results behind it.
Product listingsTurning a sparse product feed and photos into listings that can go live
| Benchmark | Claude Haiku 5.5 | Rank among available | Best available |
|---|---|---|---|
| CatalogBenchCatalogBench: reliably publish-ready | 26.8% | 14th of 20 | GPT-6 Astra 75.0% |
| CatalogBenchCatalogBench (sales brief): reliably publish-ready | 10.7% | 9th of 20 | GPT-6 Astra 69.6% |
Decks from an analysisTurning a finished analysis into a deck you could present as it is
| Benchmark | Claude Haiku 5.5 | Rank among available | Best available |
|---|---|---|---|
| DeckBenchDeckBench: deck rating | 1,130 | 10th of 19 | GPT-6 Astra 1,477 |
User surveysPlanning a user survey and reading its results without being misled
| Benchmark | Claude Haiku 5.5 | Rank among available | Best available |
|---|---|---|---|
| SurveyBenchSurveyBench: analysis rating | 1,412 | 3rd of 8 | GPT-6.1 Sol 1,650 |
Marketing planningPlanning a year of ad spend without overspending
| Benchmark | Claude Haiku 5.5 | Rank among available | Best available |
|---|---|---|---|
| ROASBenchROASBench: overall score | 24.1 | 17th of 19 | GPT-6 Astra 56.6 |
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls
| Benchmark | Claude Haiku 5.5 | Rank among available | Best available |
|---|---|---|---|
| Artificial AnalysisArtificial Analysis: Terminal-Bench 4.0 · reported | 32.8% | 15th of 105 | Claude Sonnet 5.5 63.6% |
- tau2-bench: pending. Its latest results are dated 26 May 2026, before this model was released.
- Berkeley Function Calling Leaderboard (BFCL) V4: pending. Its latest results are dated 16 Dec 2025, before this model was released.
- OpenHands Index: pending. Its latest results are dated 30 Jun 2026, before this model was released.
- Microsoft STATE-Bench: pending. Its latest results are dated 29 May 2026, before this model was released.
- Vending-Bench 2: pending. Its latest results are dated 1 Oct 2026, before this model was released.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects
- APEX-Agents: pending. Its latest results are dated 1 Oct 2026, before this model was released.
- GDP.pdf: pending. Its latest results are dated 1 Oct 2026, before this model was released.
- Remote Labor Index: pending. Its latest results are dated 1 Oct 2026, before this model was released.
Reasoning and knowledgeHard questions across science, maths and general knowledge
| Benchmark | Claude Haiku 5.5 | Rank among available | Best available |
|---|---|---|---|
| Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported | 43.4 | 17th of 198 | Claude Opus 5.5 57.6 |
| Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported | 44.4% | 20th of 197 | Claude Opus 5.5 61.4% |
WritingWhat people prefer in blind comparisons, and judged writing quality
- Arena (formerly LMArena): pending. Its latest results are dated 2 Oct 2026, before this model was released.
- UGI Leaderboard: pending. Its latest results are dated 2 Oct 2026, before this model was released.
Sticking to the factsSummarising without inventing things, and factual answers
- Vectara Hallucination Leaderboard: pending. Its latest results are dated 22 Sep 2026, before this model was released.
- Arena (formerly LMArena): pending. Its latest results are dated 2 Oct 2026, before this model was released.
- SimpleQA Verified (Epoch AI): pending. Its latest results are dated 29 Sep 2026, before this model was released.
Speed under pressureGood decisions against a real clock (fast chess)
| Benchmark | Claude Haiku 5.5 | Rank among available | Best available |
|---|---|---|---|
| BulletBenchBulletBench Bullet 60s: ladder Elo | 310 | 14th of 20 | Gemini 3.8 Flash 1,068 |
Nearest alternatives
- Open weights at similar intelligenceGLM-5.3-Flash$0.24 blended · Intelligence Index 41.8
Against Claude Haiku 5.5 at $0.20 blended, Intelligence Index 43.4. "Similar intelligence" means within 4 points or better.
Against the alternatives
The models you would most likely weigh it against: the leaders of the groups above.
Scroll sideways to see every alternative.
| Claude Haiku 5.5 | Claude Opus 5.5 most intelligent available today | Qwen3.8-Max (0902) just above overall | Grok 4.7 just below overall | |
|---|---|---|---|---|
| Price per million tokens | $0.20 | $8.00 | $3.00 | $3.00 |
| Tokens a second | 235 | 97 | 37 | 64 |
| Intelligence Index | 43.4 | 57.6 | 45.4 | 46.4 |
| Spring Prompt overall | 44 | 85 | 48 | 43 |
| Rank among available models, by use case | ||||
| Product listings | 11th of 20 | 6th | 17th | 4th |
| Decks from an analysis | 10th of 19 | 5th | 11th | 9th |
| User surveys | 3rd of 8 | 2nd | – | – |
| Marketing planning | 17th of 19 | 5th | 14th | 11th |
| Agents and tool use | 20th of 164 | 2nd | 7th | 27th |
| Professional work | pending | 3rd | 18th | – |
| Reasoning and knowledge | 19th of 198 | 1st | 17th | 16th |
| Writing | pending | 4th | 16th | 61st |
| Sticking to the facts | pending | 2nd | 38th | 106th |
| Speed under pressure | 15th of 22 | – | – | – |
| Side by side → | Side by side → | Side by side → | ||
Shaded figures are better than Claude Haiku 5.5's. A dash means no result.
Where it has been measured
- BulletBench10 results
- CatalogBench30 results
- DeckBench16 results
- ROASBench9 results
- SurveyBench19 results
- Artificial Analysis35 results
- APEX-AgentsPending: latest results 1 Oct 2026
- Arena (formerly LMArena)Pending: latest results 2 Oct 2026
- Berkeley Function Calling Leaderboard (BFCL) V4Pending: latest results 16 Dec 2025
- GDP.pdfPending: latest results 1 Oct 2026
- Microsoft STATE-BenchPending: latest results 1 Jul 2026
- OpenHands IndexPending: latest results 30 Jun 2026
- Remote Labor IndexPending: latest results 1 Oct 2026
- SimpleQA Verified (Epoch AI)Pending: latest results 29 Sep 2026
- tau2-benchPending: latest results 3 Aug 2026
- UGI LeaderboardPending: latest results 2 Oct 2026
- Vectara Hallucination LeaderboardPending: latest results 22 Sep 2026
- Vending-Bench 2Pending: latest results 1 Oct 2026
Every result
119 published results from 6 sources, each in the source's own units, with the configuration that produced it.
BulletBench · measured by Spring Prompt · 10 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| BulletBench Blitz 3+2: cost per gameUS dollars, lower is better | $0.0164 | 4 of 12 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| BulletBench Blitz 3+2: games lost on time% of games, lower is better | 37.5% | 6 of 12 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| BulletBench Blitz 3+2: invalid moves% of moves, lower is better | 0.0% | 1 of 12 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| BulletBench Blitz 3+2: ladder Eloladder Elo, higher is better | 557 370–727 |
5 of 12 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| BulletBench Blitz 3+2: median move timemilliseconds, lower is better | 6.2 s | 8 of 12 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| BulletBench Bullet 60s: cost per gameUS dollars, lower is better | $0.0019 | 6 of 31 | Claude Haiku 5.5 (no reasoning) none reasoning |
8 Oct 2026 |
| BulletBench Bullet 60s: games lost on time% of games, lower is better | 37.5% | 17 of 31 | Claude Haiku 5.5 (no reasoning) none reasoning |
8 Oct 2026 |
| BulletBench Bullet 60s: invalid moves% of moves, lower is better | 0.0% | 1 of 31 | Claude Haiku 5.5 (no reasoning) none reasoning |
8 Oct 2026 |
| BulletBench Bullet 60s: ladder Eloladder Elo, higher is better | 310 0–575 |
23 of 31 | Claude Haiku 5.5 (no reasoning) none reasoning |
8 Oct 2026 |
| BulletBench Bullet 60s: median move timemilliseconds, lower is better | 1.7 s | 23 of 31 | Claude Haiku 5.5 (no reasoning) none reasoning |
8 Oct 2026 |
Not shown: Not your cost: prices are those charged on the run date.
CatalogBench · measured by Spring Prompt · 30 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| CatalogBench (sales brief): channel rules broken% of products, lower is better | 0.0% | 1 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): claims to check% of products, lower is better | 1.2% | 12 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): content quality% of checks, higher is better | 96.3% | 13 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): failed outputs% of products, lower is better | 0.0% | 1 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): missing UK information% of products, lower is better | 0.0% | 1 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): not findable% of products, lower is better | 10.7% | 20 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): publish-ready listings% of products, higher is better | 20.2% | 11 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): reliably publish-ready% of products, higher is better | 10.7% | 9 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): unsupported claims% of products, lower is better | 70.2% | 12 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): unsupported claimsclaims per product, lower is better | 2.46 | 12 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): wrong attributes% of products, lower is better | 18.5% | 10 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench (sales brief): wrong category or variant% of products, lower is better | 0.0% | 1 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: channel compliance% of products, higher is better | 100.0% | 1 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: channel rules broken% of products, lower is better | 0.6% | 11 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: claims to check% of products, lower is better | 0.6% | 13 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: conflicts caught% of conflicts, higher is better | 100.0% | 1 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: content quality% of checks, higher is better | 95.6% | 15 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: cost per productUS dollars, lower is better | $0.0012 | 2 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: decision accuracy% of decisions, higher is better | 95.6% | 14 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: failed outputs% of products, lower is better | 0.0% | 1 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: field accuracy% of missing fields, higher is better | 91.0% | 18 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: invented values% of filled values, lower is better | 0.8% | 8 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: missing UK information% of products, lower is better | 0.0% | 1 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: not findable% of products, lower is better | 11.9% | 19 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: publish-ready listings% of products, higher is better | 47.0% | 13 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: reliably publish-ready% of products, higher is better | 26.8% | 14 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: unsupported claims% of products, lower is better | 32.1% | 13 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: unsupported claimsclaims per product, lower is better | 1.20 | 16 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: wrong attributes% of products, lower is better | 19.6% | 13 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| CatalogBench: wrong category or variant% of products, lower is better | 0.0% | 1 of 21 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
Not shown: Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
DeckBench · measured by Spring Prompt · 16 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| DeckBench: accurate decks% of tasks, higher is better | 33.3% | 12 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: caveat dropped% of tasks, lower is better | 16.7% | 10 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: clean layout% of tasks, higher is better | 66.7% | 6 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: cost per deckUS dollars, lower is better | $0.0091 | 2 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: deck ratingrating, higher is better | 1,130 995–1,258 |
10 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: design quality% of the maximum, higher is better | 61.1% | 12 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: draft figure quoted% of tasks, lower is better | 33.3% | 11 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: findings missing% of tasks, lower is better | 16.7% | 14 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: head-to-head win rate% of comparisons, higher is better | 50.0% | 11 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: layout defects% of tasks, lower is better | 33.3% | 6 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: misleading metric used% of tasks, lower is better | 16.7% | 19 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: presentable decks% of tasks, higher is better | 33.3% | 5 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: recommendation late or wrong% of tasks, lower is better | 33.3% | 10 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: slides needing work% of tasks, lower is better | 16.7% | 2 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: unsupported claims% of tasks, lower is better | 33.3% | 13 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| DeckBench: unsupported numbers% of tasks, lower is better | 16.7% | 17 of 19 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
Not shown: Not your analysis or your brand: invented companies, six public tasks, one deck per model per task at the provider's default reasoning setting.
ROASBench · measured by Spring Prompt · 9 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| ROASBench: audience scorescore out of 100, higher is better | 48.4 | 15 of 20 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| ROASBench: business scorescore out of 100, higher is better | 13.5 | 16 of 20 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| ROASBench: consistency scorescore out of 100, higher is better | 10.6 | 16 of 20 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| ROASBench: contribution profitsimulated US dollars, higher is better | $32,146 | 16 of 20 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| ROASBench: cost of a runUS dollars, lower is better | $0.0313 | 3 of 20 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| ROASBench: months over budgetmonths of 12, lower is better | 0 | 1 of 20 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| ROASBench: overall scorescore out of 100, higher is better | 24.1 | 18 of 20 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| ROASBench: planning scorescore out of 100, higher is better | 54.8 | 11 of 20 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| ROASBench: return on ad spendprofit per $1 spent, higher is better | $0.03 | 16 of 20 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
Not shown: Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
SurveyBench · measured by Spring Prompt · 19 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| SurveyBench: analysis ratingrating, higher is better | 1,412 1,322–1,492 |
3 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: analysis scorepoints of 100, higher is better | 88.4 | 3 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: cost per taskUS dollars, lower is better | $0.0185 | 2 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: head-to-head win rate% of comparisons, higher is better | 72.5% | 3 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: leading or double-barrelled questions% of tasks, lower is better | 33.3% | 3 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: mistaken next steps% of tasks, lower is better | 22.2% | 3 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: numbers right% of answers, higher is better | 100.0% | 1 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: objectives unanswerable% of tasks, lower is better | 0.0% | 1 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: precision ignored% of tasks, lower is better | 0.0% | 1 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: research scorepoints of 100, higher is better | 90.2 | 3 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: sound research% of tasks, higher is better | 11.1% | 3 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: stakeholder's view presumed% of tasks, lower is better | 0.0% | 1 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: survey plan scorepoints of 100, higher is better | 94.5 | 3 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: survey structure% of tasks, lower is better | 22.2% | 6 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: traps handled% of traps, higher is better | 94.2% | 3 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: traps missed% of tasks, lower is better | 22.2% | 1 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: unsupported findings% of tasks, lower is better | 77.8% | 4 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: wrong numbers% of tasks, lower is better | 0.0% | 1 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
| SurveyBench: wrong recommendation% of tasks, lower is better | 11.1% | 2 of 8 | Claude Haiku 5.5 provider default reasoning |
8 Oct 2026 |
Not shown: Not your customers or your data: invented organisations with simulated respondents, nine public tasks, one run per model per task at the provider's default reasoning setting.
Artificial Analysis · reported by Artificial Analysis · 35 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better | 37.8 | 75 of 407 | Claude Haiku 5.5 (high reasoning) | 8 Oct 2026 |
| Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better | 29.4 | 117 of 407 | Claude Haiku 5.5 (low reasoning) | 8 Oct 2026 |
| Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better | 43.4 | 42 of 407 | Claude Haiku 5.5 (max reasoning) | 8 Oct 2026 |
| Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better | 34.5 | 85 of 407 | Claude Haiku 5.5 (medium reasoning) | 8 Oct 2026 |
| Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better | 41.2 | 52 of 407 | Claude Haiku 5.5 (xhigh reasoning) | 8 Oct 2026 |
| Artificial Analysis: Humanity's Last Exam% of questions, higher is better | 37.3% | 97 of 405 | Claude Haiku 5.5 (high reasoning) | 8 Oct 2026 |
| Artificial Analysis: Humanity's Last Exam% of questions, higher is better | 27.0% | 160 of 405 | Claude Haiku 5.5 (low reasoning) | 8 Oct 2026 |
| Artificial Analysis: Humanity's Last Exam% of questions, higher is better | 44.4% | 47 of 405 | Claude Haiku 5.5 (max reasoning) | 8 Oct 2026 |
| Artificial Analysis: Humanity's Last Exam% of questions, higher is better | 33.8% | 119 of 405 | Claude Haiku 5.5 (medium reasoning) | 8 Oct 2026 |
| Artificial Analysis: Humanity's Last Exam% of questions, higher is better | 42.7% | 57 of 405 | Claude Haiku 5.5 (xhigh reasoning) | 8 Oct 2026 |
| Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better | 77.3% | 117 of 393 | Claude Haiku 5.5 (high reasoning) | 8 Oct 2026 |
| Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better | 72.0% | 172 of 393 | Claude Haiku 5.5 (low reasoning) | 8 Oct 2026 |
| Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better | 82.7% | 30 of 393 | Claude Haiku 5.5 (max reasoning) | 8 Oct 2026 |
| Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better | 77.3% | 117 of 393 | Claude Haiku 5.5 (medium reasoning) | 8 Oct 2026 |
| Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better | 78.3% | 100 of 393 | Claude Haiku 5.5 (xhigh reasoning) | 8 Oct 2026 |
| Artificial Analysis: Output speedtokens per second, higher is better | 143 | 34 of 133 | Claude Haiku 5.5 (high reasoning) | 8 Oct 2026 |
| Artificial Analysis: Output speedtokens per second, higher is better | 162 | 25 of 133 | Claude Haiku 5.5 (low reasoning) | 8 Oct 2026 |
| Artificial Analysis: Output speedtokens per second, higher is better | 235 | 5 of 133 | Claude Haiku 5.5 (max reasoning) | 8 Oct 2026 |
| Artificial Analysis: Output speedtokens per second, higher is better | 147 | 30 of 133 | Claude Haiku 5.5 (medium reasoning) | 8 Oct 2026 |
| Artificial Analysis: Output speedtokens per second, higher is better | 203 | 16 of 133 | Claude Haiku 5.5 (xhigh reasoning) | 8 Oct 2026 |
| Artificial Analysis: SciCode% of problems, higher is better | 48.7% | 114 of 191 | Claude Haiku 5.5 (high reasoning) | 8 Oct 2026 |
| Artificial Analysis: SciCode% of problems, higher is better | 49.2% | 110 of 191 | Claude Haiku 5.5 (low reasoning) | 8 Oct 2026 |
| Artificial Analysis: SciCode% of problems, higher is better | 55.0% | 49 of 191 | Claude Haiku 5.5 (max reasoning) | 8 Oct 2026 |
| Artificial Analysis: SciCode% of problems, higher is better | 49.0% | 113 of 191 | Claude Haiku 5.5 (medium reasoning) | 8 Oct 2026 |
| Artificial Analysis: SciCode% of problems, higher is better | 51.7% | 83 of 191 | Claude Haiku 5.5 (xhigh reasoning) | 8 Oct 2026 |
| Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better | 21.7% | 52 of 187 | Claude Haiku 5.5 (high reasoning) | 8 Oct 2026 |
| Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better | 12.6% | 73 of 187 | Claude Haiku 5.5 (low reasoning) | 8 Oct 2026 |
| Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better | 32.8% | 36 of 187 | Claude Haiku 5.5 (max reasoning) | 8 Oct 2026 |
| Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better | 15.2% | 64 of 187 | Claude Haiku 5.5 (medium reasoning) | 8 Oct 2026 |
| Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better | 29.3% | 42 of 187 | Claude Haiku 5.5 (xhigh reasoning) | 8 Oct 2026 |
| Artificial Analysis: Time to first answer tokenseconds, lower is better | 18.8 | 83 of 133 | Claude Haiku 5.5 (high reasoning) | 8 Oct 2026 |
| Artificial Analysis: Time to first answer tokenseconds, lower is better | 5.08 | 50 of 133 | Claude Haiku 5.5 (low reasoning) | 8 Oct 2026 |
| Artificial Analysis: Time to first answer tokenseconds, lower is better | 334 | 132 of 133 | Claude Haiku 5.5 (max reasoning) | 8 Oct 2026 |
| Artificial Analysis: Time to first answer tokenseconds, lower is better | 8.09 | 58 of 133 | Claude Haiku 5.5 (medium reasoning) | 8 Oct 2026 |
| Artificial Analysis: Time to first answer tokenseconds, lower is better | 63.5 | 120 of 133 | Claude Haiku 5.5 (xhigh reasoning) | 8 Oct 2026 |
Not shown: Not business work, and a blend: read the parts for any one task.