Anthropic
Claude Fable 5
- Benchmark records
- 11
- Snapshot
- 2026-07-30
- Release date
- 2026-06-09
Independent model comparison
Compare the models using only evidence that can be identified and attributed. Exact benchmark matches appear first; independent facts and user reports remain separate and do not create an overall ranking.
Exact comparison coverage
8 benchmark records share an exact reviewed comparison key.
Evidence strength
Editorial review pending
Substantial matched evidence exists, but this pair has not passed the pair-native publication gate.
Models at a glance
Anthropic
Anthropic
| Detail | Claude Fable 5 | Claude Sonnet 4.6 |
|---|---|---|
| Provider |
Anthropic |
Anthropic |
| Release date Shown only where an exact reviewed release date is available. |
2026-06-09 |
2026-02-17 |
| Input price Artificial Analysis operational reference; see the attributed details below. |
$10.00 / 1M tokens |
$3.00 / 1M tokens |
| Output price Artificial Analysis operational reference; see the attributed details below. |
$50.00 / 1M tokens |
$15.00 / 1M tokens |
| Median output speed Artificial Analysis operational reference; see the attributed details below. |
65.0 tok/s |
46.9 tok/s |
| Exact benchmark coverage Only records carrying the same reviewed comparison identity are counted. |
8 shared records |
8 shared records |
Price and speed appear only when an exact reviewed operational reference is attached to that model identity.
Primary comparison evidence
A shared benchmark name is not enough. A result appears side by side only when both model records carry the same reviewed comparison key. Duplicate or unversioned records are omitted rather than guessed into alignment.
A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.
Claude Fable 5
60
Artificial Analysis Intelligence Index
Configuration: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30
Claude Sonnet 4.6
34*
Artificial Analysis Intelligence Index
Configuration: Claude Sonnet 4.6 (Non-reasoning, Low Effort)
Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30
The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.
Claude Fable 5
93%
GPQA Diamond
Configuration: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Claude Sonnet 4.6
80%
GPQA Diamond
Configuration: Claude Sonnet 4.6 (Non-reasoning, Low Effort)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A broad expert-level benchmark of difficult academic reasoning and knowledge questions.
Claude Fable 5
53%
Humanity's Last Exam
Configuration: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Claude Sonnet 4.6
11%
Humanity's Last Exam
Configuration: Claude Sonnet 4.6 (Non-reasoning, Low Effort)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.
Claude Fable 5
63%
IFBench
Configuration: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Claude Sonnet 4.6
42%
IFBench
Configuration: Claude Sonnet 4.6 (Non-reasoning, Low Effort)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.
Claude Fable 5
95.8%
OpenHands SWE-Bench resolved
Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.
Configuration: OpenHands v1.28.0 + claude-fable-5
OpenHands Index · as of 30 Jun 2026 · OpenHands · SWE-Bench 2026.06.30-3015ac6
Claude Sonnet 4.6
74.4%
OpenHands SWE-Bench resolved
Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.
Configuration: OpenHands v1.11.5 + claude-sonnet-4-6
OpenHands Index · as of 30 Jun 2026 · OpenHands · SWE-Bench 2026.06.30-3015ac6
A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.
Claude Fable 5
3 configurations
Average ROASBench score
Configuration: 3 configurations compared separately
Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation
Claude Sonnet 4.6
21.21
Average ROASBench score
Configuration: Anthropic: Claude Sonnet 4.6
Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation
Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.
Claude Fable 5
85.09%
Direct benchmark score
Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.
Configuration: claude-fable-5
Structured Output Benchmark (SOB) · as of 17 Jul 2026 · Sob-v1@da785a8521c8954283b2989d01e54d80c4e023c6:upstream-provider-configs:temperature-0-where-supported:max-output-2048:reasoning-disabled-or-minimum-where-required:official-modality-weights
Claude Sonnet 4.6
85.44%
Direct benchmark score
Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.
Configuration: Claude-Sonnet-4.6
Structured Output Benchmark (SOB) · as of 17 Jul 2026 · Sob-v1@da785a8521c8954283b2989d01e54d80c4e023c6:upstream-provider-configs:temperature-0-where-supported:max-output-2048:reasoning-disabled-or-minimum-where-required:official-modality-weights
A legacy agentic tool-use benchmark for conversational work in a telecom environment.
Claude Fable 5
99%
τ²-Bench Telecom
Configuration: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Claude Sonnet 4.6
79%
τ²-Bench Telecom
Configuration: Claude Sonnet 4.6 (Non-reasoning, Low Effort)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Secondary context · not head-to-head
The sources below evaluated or discussed each model independently. Putting them in adjacent columns makes them easier to inspect, but it does not make their metrics, samples, or observations directly comparable.
Attributed third-party facts
These facts retain their source, date, protocol, and model identity. They are contextual model-level signals—not a Spring Prompt overall score.
Input price
$10.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Output price
$50.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Median output speed
65.0 tok/s
Median output tokens received per second after generation begins; this excludes time to first token and is not end-to-end response latency.
Artificial Analysis · as of 16 Jul 2026
Input price
$3.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Output price
$15.00 / 1M tokens
Representative configuration on the cited source page.
Artificial Analysis · as of 16 Jul 2026
Median output speed
46.9 tok/s
Median output tokens received per second after generation begins; this excludes time to first token and is not end-to-end response latency.
Artificial Analysis · as of 16 Jul 2026
Anecdotal model-level reports
Recurring themes from each model's declared observation window. They are not a representative survey and are not direct comparisons between these models.
Initial community opinions
In the search-indexed launch-window sample, Claude Fable 5 appeared especially capable on difficult code review, architecture, and long-running agent work when tasks were tightly scoped. Reports also repeatedly noted slow interactive responses, high usage cost, and safeguards or product routing that could interrupt the expected model experience, so early results varied substantially by workflow.
Source venue: Reddit · window: 2026-06-09 to 2026-06-23 · observed items: 4
Initial community opinions
The search-indexed launch-window sample highlighted Claude Sonnet 4.6's speed, price, task focus, and occasional strong coding output. Recurring complaints involved superficial reasoning, distraction, instruction failures, hallucinations, and a less conversational style. Reports varied enough that early usefulness appeared closely tied to task size and the surrounding product.
Source venue: Reddit · window: 2026-02-17 to 2026-03-03 · observed items: 4
Independent early-test reports
Editor-paraphrased reports from independent authors in each model's fixed 14-day launch window. These selectively surfaced anecdotes are not direct head-to-head tests unless they also appear in the dedicated direct-report section above.
Early technical field tests
Initial field demonstrations showed Fable 5 turning personal sports footage into a tailored visual-analysis overlay and building a navigable, real-scale Yosemite scene from geospatial data. Both artifacts point to strong prompt-to-product work across visual reasoning and procedural 3D, but they are self-selected demonstrations and do not establish routine reliability or failure rates.
Source venue: X · window: 2026-06-09 to 2026-06-23 · 2 tests · 2 independent authors · 2 with a method or artifact · identity reviewed
Early technical field tests
Early Sonnet 4.6 tests covered conditional routing, scheduling, contract workflows, and repeated Claude Code use. Workflow automation looked capable, but a practical coding comparison found that longer reasoning and more tool calls could slow end-to-end completion despite respectable token speed. The launch evidence supports broad utility with a real throughput tradeoff.
Source venue: X · window: 2026-02-17 to 2026-03-03 · 2 tests · 2 independent authors · 2 with a method or artifact · identity reviewed
X field tests remain separate from Reddit community opinions, Artificial Analysis facts, and every scored benchmark.
Spring Prompt does not name an overall winner from unrelated benchmark scores, third-party metrics, Reddit opinions, or X field-test reports.
Different benchmarks measure different constructs and may use different populations, prompts, harnesses, model revisions, tools, and score scales. This page does not average those values or infer a winner from source coverage. Test both models on your own production work before making a consequential choice.