Claude Haiku 5.5 Is Far Better Than Haiku 4.5 at a Tenth of the Price. GPT-6 Luna Still Beats It at Business Work.

Anthropic released Claude Haiku 5.5 on 7 October. It is the first Haiku that thinks before it answers. It has a 1-million-token context window, and it costs a tenth of what Haiku 4.5 did.
We ran it through our business benchmarks and all 18 playground tests the next day. It is a big step up from Haiku 4.5 on everything we measure. But it is up against GPT-6 Luna, which costs the same per token. Luna beat it on product listings, decks and ad planning, and used fewer tokens doing it.
Claude Haiku 5.5 at a glance
- Much better than Haiku 4.5. Reliably publish-ready product listings went from 1.8% to 26.8%, and its ad-planning year went from a loss to a small profit.
- Behind GPT-6 Luna at the same price. Luna's listings were twice as often ready to publish (55.4%), and it scored more than twice as high at ad planning (53.4 against 24.1).
- Cheap, but wordy. Its default setting thinks at medium effort, so a CatalogBench product cost $0.0012, against $0.0005 for Luna.
- Strong and very cheap on small coding tasks. In our playground it passed every check on 10 of 18 tests, matching our pick on those, for $0.11 for all 18 answers.
- Not for split-second decisions. Even with thinking off, its moves took 1.7 seconds at the median, and it was too slow for our 10-second chess clock.
- Launch-day reaction was warm: 78% of the posts and comments that took a side were positive, mostly about the price cut. The critical ones were hands-on tests in which GPT-6 Luna did the same job faster or for fewer tokens.
- Price: $0.10 in and $0.50 out per million tokens for prompts up to 100,000 tokens; $0.50 and $2.50 above that.
Specs, price and what Anthropic claims
According to Anthropic's announcement, its documentation and the system card:
- Price: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Above that, the whole request costs $0.50 and $2.50. It is the only current Claude model that charges more for long prompts. Anthropic also says its new tokenizer turns the same text into about 30% more tokens than Haiku 4.5's, so it puts the saving at "around 75%" rather than 90%.
- Context and output: 1 million tokens of context and up to 128,000 tokens of output. Text and images in, text out. Knowledge cutoff June 2026.
- Thinking: adaptive thinking is on by default at medium effort, with five settings from low to max. You can turn thinking off at high effort or below.
- Claimed strengths: high-volume work such as summaries, classification and extraction, sub-agents for Opus and Sonnet, customer support and browser use. Anthropic says Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding tasks".
- Headline numbers, all at max effort: 72.4% on OSWorld 2.1, 39.2% on Terminal-Bench 4.0, 45.9% on Humanity's Last Exam without tools, and 1620 on GDPval-AA, ahead of GPT-6 Luna on every benchmark in its table.
Note the setting. Every headline number is at max effort, but the API defaults to medium. The system card itself shows the gap: on GDPval-AA, Haiku 5.5 scores 1620 at max and 1277 at medium, using "about a tenth of the output tokens".
Claude Haiku 5.5 benchmarks, measured by Spring Prompt
As on our model pages, we rank it only among models available today (those you can call through OpenRouter), and within the groups it belongs to:
- Same price: under $0.25 per million tokens, blending three input tokens to one output. Haiku 5.5 costs $0.20. Of the models we have run, only GPT-6 Luna is in this band, so we compare the two directly below.
- Similar intelligence: within 4 points of its 43.4 on the Artificial Analysis Intelligence Index. That score is at max effort; at the default medium setting we ran, Artificial Analysis gives it 34.5.
- Same speed: 200 or more tokens a second. Artificial Analysis measures 235 at max effort.
How to read this table: each row is one of our benchmarks. Reading across, it shows where Claude Haiku 5.5 ranks among every model available today, against GPT-6 Luna at the same price, and among only the models of similar intelligence and speed. Under each rank is the best model in that group and its score.
| Benchmarkwhat it measures | Claude Haiku 5.5its score | Rank among all modelsavailable today | Against GPT-6 Lunathe same price, $0.20 | Rank at its intelligencewithin 4 points on the AA index | Rank at its speed200 or more tokens a second |
|---|---|---|---|---|---|
| CatalogBenchshare of product listings reliably ready to publish | 26.8%default setting (medium) | 14th of 20Best: GPT-6 Astra, 75.0% | BehindGPT-6 Luna, 55.4% | 6th of 6 (last)Best: Grok 4.7, 57.1% | 2nd of 3Best: DeepSeek V4.1 Flash, 41.1% |
| CatalogBench sales briefshare of sales briefs reliably ready to publish | 10.7%default setting (medium) | 9th of 20Best: GPT-6 Astra, 69.6% | BehindGPT-6 Luna, 51.8% | 3rd of 6Best: Grok 4.7, 33.9% | 2nd of 3Best: DeepSeek V4.1 Flash, 23.2% |
| DeckBenchdeck rating from head-to-head comparisons | 1,130default setting (medium) | 10th of 19Best: GPT-6 Astra, 1,477 | BehindGPT-6 Luna, 1,259 | 3rd of 5Best: Kimi K3, 1,229 | Not rankedonly 2 of these models measured |
| ROASBenchscore out of 100 for a year of ad-spend planning | 24.1default setting (medium) | 17th of 19Best: GPT-6 Astra, 56.6 | BehindGPT-6 Luna, 53.4 | 6th of 6 (last)Best: Gemini 3.8 Flash, 51.4 | Not rankedonly 2 of these models measured |
| BulletBench Bulletchess rating with 60 seconds a game | 310 Elowith thinking off; too slow at medium and low | 14th of 20Best: Gemini 3.8 Flash, 1,068 | AheadGPT-6 Luna, 249 | 2nd of 4Best: Gemini 3.8 Flash, 1,068 | Not rankedonly 2 of these models measured |
| SurveyBenchanalysis rating from head-to-head comparisons | 1,412default setting (medium) | 3rd of 8Best: GPT-6.1 Sol, 1,650 | Not measuredGPT-6 Luna has not taken SurveyBench | Not rankedno other model in this group measured yet | Not rankedonly 2 of these models measured |
Each model is ranked at its better reasoning setting. A group only counts models that benchmark has measured, and a group of fewer than three is not ranked. "Best" is the top model in that group, not necessarily one you would choose instead: see the alternatives.
Against Haiku 4.5, the gains are large on every benchmark. Against the models you would actually weigh it against, it is mid-table at best. GPT-6 Luna, at the same price, beat it on listings, sales briefs, decks and ad planning. Among models of similar intelligence, it is last on listings and on ad planning.
Product listings: fewer inventions than Haiku 4.5, still many claims
CatalogBench gives each model a sparse product feed and product photos, and asks for listings that could go live. Haiku 5.5 rarely invents product attributes: 0.8% of the fields it filled were made up, the lowest of the three models here (GPT-6 Luna 1.7%, Haiku 4.5 6.2%). It caught every conflict between feed and image and never failed to return an output.
Its weakness is the copy. 60.1% of its listings contained at least one claim the product data did not support, an average of 1.2 per product. That is a big improvement on Haiku 4.5 (83.9%, 2.7 per product), but GPT-6 Luna's figure was 12.5%. Asked for "compelling, SEO-optimised copy that sells", Haiku 5.5's share rose to 78.0%. Only 26.8% of its listings were publish-ready in all three runs, against 55.4% for Luna.
Decks: a middling rating, but presentable more often than Luna
On DeckBench, where models turn a finished analysis into a deck, Haiku 5.5 rated 1,130, 10th of 19 and well above Haiku 4.5 (688). Three of its eight decks were presentable as they were, which GPT-6 Luna managed for none. But the judges still preferred Luna's decks overall (1,259), and Haiku 5.5 returned no deck at all for one task.
Ad-spend planning: profitable, just
On ROASBench, a 12-month simulated ad-budget game, Haiku 5.5 scored 24.1, up from 17.5 for Haiku 4.5. It made a simulated $32,146 in contribution profit, where Haiku 4.5 lost $205,001, and it never went over budget. But it lost money in five of the twelve months, and GPT-6 Luna made $738,879 for a score of 53.4.
User surveys: its best result
SurveyBench asks a model to plan a user survey and then read its results without falling for planted traps, such as a filtered base or a sample of current customers only. It is our newest benchmark, with eight models so far. Haiku 5.5 rated 1,412, third behind GPT-6.1 Sol (1,650) and Claude Opus 5.5 (1,646), and far ahead of Mistral Large 4 (815), Gemini 3.1 Pro (810) and Haiku 4.5 (465). Across all 12 tasks it gave every number correctly and handled 96% of the traps, for $0.23 in total. Its survey plans were sound for 7 of 12 tasks, but its analyses for only 3: most of the rest drew at least one finding the responses did not support.
Speed: fast to write, slow to start
Artificial Analysis measures Haiku 5.5 writing 235 tokens a second at max effort and 137 at medium, which is quick. Chess shows a different side. With three minutes a game it made a typical move in 6.2 seconds and reached 557 ladder Elo, above Haiku 4.5's 491. But at the 10-second Lightning clock it lost its first four games on time at every setting we tried, including with thinking off, and at the 60-second Bullet clock it only finished games with thinking off (310 Elo). Even then, a typical move took 1.7 seconds, against 1.3 for Haiku 4.5. These runs were on launch day, when providers are busiest, so we will re-run them.
In the playground: 18 tests for 11 cents
We also gave Haiku 5.5 every playground test, the same short visual and interactive tasks we gave ten other models: games, simulations, drawings, landing pages and business pages, each a single self-contained file. We played every answer and marked each test's checks by hand.
All 18 answers together cost $0.11. The next cheapest model, DeepSeek V4.1 Flash, cost $0.30 for the same tests, and Claude Opus 5.5 cost $6.09. Haiku 5.5 was the cheapest answer on 16 of the 18 tests and the fastest on 12.
On 10 tests it passed every check, as our pick did, though our picks were more polished each time:
- Chess with every rule: castling, en passant, promotion, checkmate and stalemate all worked, for $0.011 against $0.075 for GPT-6.1 Sol.
- A product page from one feed row: every line traced back to the feed, with no invented claims, for $0.0033 against $0.37 for Claude Opus 5.5.
- Conway's Game of Life, a category page, a pricing table, an analogue clock, Flappy Bird, the solar system, a ball in a spinning hexagon and an animated logo also passed every check.
It struggled where a test needs physics or careful drawing:
- Newton's cradle: the balls push each other apart at a distance, so the swing never transfers.
- The jelly effect: a ball that slides about rather than a jelly that wobbles.
- A map of the world: Africa's northern coast is drawn south of the Equator.
- A newsletter email: it invented product facts, including "made in small batches in the UK".
Every answer, with our notes, is on Claude Haiku 5.5's playground page.
Third-party benchmarks
- Artificial Analysis (reported 7 October): Intelligence Index 43.4 at max effort, 41.2 at xhigh, 37.8 at high and 34.5 at the default medium setting, against 16.9 for Haiku 4.5 and 38.1 for GPT-6 Luna. Artificial Analysis calls it "notably fast, however very verbose": it used about 162,000 output tokens per Intelligence Index task at max effort, about three times GPT-6 Luna. On its AA-Omniscience test it answered fewer questions correctly than Luna or Gemini 3.8 Flash, but made things up less often (a 40% hallucination rate against 77% and 55%). Artificial Analysis says two of its figures are provisional: its AutomationBench-AA score, which a pre-release over-refusal problem probably lowered, and its cost figures, which do not yet include the higher price above 100,000 tokens.
- Pending: Arena, tau2-bench, BFCL, OpenHands, STATE-Bench, Vending-Bench 2, the Vectara Hallucination Leaderboard, SimpleQA Verified and the professional-work sources (APEX-Agents, GDP.pdf, Remote Labor Index).
What people are saying about Claude Haiku 5.5
Reactions from launch day, on X, blogs and Hacker News. Quotes are as posted; each card links to the original.
How launch day read: 78% positive, 9% mixed, 13% negative of the 99 posts and comments that took a side
NegativeMixedPositive
On 8 October we read the top posts on X for "Haiku 5.5" and "Claude Haiku 5.5", and every top-level comment on the Hacker News launch thread. We labelled each one positive, mixed, negative or neutral. News, jokes, questions, and posts from Anthropic, its staff and its launch partners count as neutral or are left out of the bars. Labels are by Claude Opus 5.5 against a written rubric, in a single pass; we have not yet checked a sample by hand. X ranks top posts by engagement, so its sample leans to the loudest voices, and on X most posts repeated Anthropic's launch claims. Every labelled post.
On X the reaction was overwhelmingly warm, mostly about the price cut and Anthropic's benchmark charts. The most-shared critical posts were hands-on tests, and they made the same point as our results: at the same price, GPT-6 Luna often did the job faster and for fewer tokens. A selection, with how we read each one:
Anthropic has released Claude Haiku 5.5, scoring 43 on the Artificial Analysis Intelligence Index - up 26 points one year after the last Haiku release
Anthropic somehow made Haiku MUCH smarter... then made it ~75% cheaper to run. […] Anthropic's cheap, fast model literally beats its old flagship on BOTH knowledge-work benchmarks.
Haiku 5.5 in the wild on SaaStr AI apps (out today):
- 1.5x faster
- 85% cheaper
- real reasoning (much more useable) […] Not all AI is inflationary
If your workloads fit in 100,000 tokens, Haiku is the same price as Luna and reports higher benchmark scores. Above 100,000 tokens, Luna looks like a much better deal.
9x cheaper than Haiku 4.5 and 2 letter grades better. […] Compared to OpenAI: GPT-6 Luna did a bit better and was about 30% the cost of Haiku 5.5.
I was curious if this smaller model was good enough for complex UI. It was not. […] Very fast though, and likely best used for small subagent tasks / tightly scoped work.
WTF is going on with Haiku 5.5 at max reasoning effort it's more token-hungry than opus 5.5 max and uses over 2x more than fable 5.1 max
Claude Haiku 5.5 just dropped and it completely fails the BridgeBench lava lamp test. […] GPT 6 Luna is over 6x faster and 3x cheaper.
Fans welcomed a Haiku that matches GPT-6 Luna's price and beats it on Anthropic's charts, and early testers praised its speed. Critics said the price step at 100,000 tokens comes too soon, that it burns tokens, and that their first tests were weak. Many comments were about the new API credits for subscribers instead.
Claude Haiku 5.5 vs GPT-6 Luna, Gemini 3.8 Flash and DeepSeek V4.1 Flash
| Model | Why compare | Price* | Tokens a second | AA Intelligence | Spring Prompt overall |
|---|---|---|---|---|---|
| Claude Haiku 5.5 | — | $0.20 | 235 | 43.4 | 44 |
| GPT-6 Luna compare | The same price | $0.20 | 127 | 38.1 | 41 |
| DeepSeek V4.1 Flash compare | Similar intelligence and speed | $0.52 | 217 | 39.5 | 58 |
| Gemini 3.8 Flash compare | Similar intelligence | $1.50 | 147 | 40.9 | 61 |
| Claude Haiku 4.5 compare | The previous Haiku | $2.00 | 97 | 16.9 | 8 |
| Claude Sonnet 5.5 compare | Anthropic's next step up | $4.00 | 134 | 56.0 | 64 |
*List price per million tokens, blended at three input tokens to one output, on 8 October, for prompts up to 100,000 tokens. Speed and intelligence are from Artificial Analysis, at each model's best setting: for Haiku 5.5, max effort.
For most business text, GPT-6 Luna is the better buy at the same price: it invented fewer claims, planned better and used fewer tokens. DeepSeek V4.1 Flash and Gemini 3.8 Flash cost more but score well above both on our overall rating. If you are moving up from Haiku 4.5, Haiku 5.5 is better on everything we measure and costs a tenth as much, but test it against Luna before you settle.
Is Claude Haiku 5.5 good? Who should use it, and for what
- Teams already on Haiku 4.5. It is better on every benchmark we run, and much cheaper per token.
- Small, self-contained coding and front-end tasks. It passed every check on 10 of our 18 playground tests for under a penny each.
- Extracting structured data from images and feeds. It rarely invented a product attribute and caught every conflict between feed and photo.
- Not for customer-facing copy without a fact check. 60% of its product listings contained a claim the data did not support.
- Not where every answer must arrive within a second or two. Even with thinking off, its typical reply took 1.7 seconds in our chess runs.
- Watch long prompts. Above 100,000 tokens the price rises fivefold, which changes the sums for long documents.
As always, these benchmarks aren't your data, so test it on your own prompts before switching. Try it against GPT-6 Luna side by side in the playground, and if it would write your catalogue, our catalogue feed diagnostic measures how often it invents attributes on your real feed.
What we'll update
We'll publish a week-one verdict once Arena, tau2-bench and the hallucination leaderboards have results, and re-run BulletBench away from launch-day traffic. We may also run our benchmarks at high effort, which Anthropic recommends for knowledge work. Until then, the Claude Haiku 5.5 model page has every result as it lands.
Quick answers
How much does Claude Haiku 5.5 cost?
Anthropic's list price is $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. A longer prompt is billed at $0.50 and $2.50 for the whole request. That is a tenth of Claude Haiku 4.5's price for short prompts and half for long ones.
What is Claude Haiku 5.5's context window?
1 million tokens, with up to 128,000 tokens of output (300,000 on the Message Batches API with a beta header). Haiku 4.5 had 200,000 tokens.
Does Claude Haiku 5.5 think before it answers?
Yes, by default. Adaptive thinking is on at medium effort, and there are five effort settings from low to max. Anthropic's headline benchmark scores are at max effort, not the default.
How good is Claude Haiku 5.5?
Artificial Analysis gives it 43.4 on its Intelligence Index at max effort and 34.5 at the default medium setting, against 16.9 for Haiku 4.5. On Spring Prompt's business benchmarks it is mid-table among models available today, and behind GPT-6 Luna, which costs the same.
Does Claude Haiku 5.5 make things up?
Less than Haiku 4.5, but still often. On CatalogBench, 60% of its product listings contained at least one claim the product data did not support, against 84% for Haiku 4.5 and 12.5% for GPT-6 Luna.
Claude Haiku 5.5 or GPT-6 Luna?
They cost the same per token. GPT-6 Luna beat Haiku 5.5 on our product listing, deck and ad-planning benchmarks and used fewer tokens per task. Haiku 5.5 writes faster and scores higher on the Intelligence Index at max effort.