Provider
Anthropic
Base model family
Model intelligence profile
This family profile brings together 13 published benchmark records for Anthropic Claude Fable 5. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.
Provider
Anthropic
Base model family
Release date
2026-06-09
Exact-family source only
Published benchmarks
13
Native scales kept separate
Artificial Analysis
Reference available
As of 2026-07-16
Benchmark-level evidence
These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.
Spring Prompt benchmark
79
60-second bullet Elo
Chess decision quality under a shared 60-second game clock and a 10+1 per-move clock.
Ladder Elo: anchored to a Stockfish skill ladder (random mover = 400). Internally consistent ordering, not FIDE-calibrated.
Claude Fable 5 (low)
79
Spring Prompt benchmark
3 configurations
Average ROASBench score
A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.
Anthropic: Claude Fable 5
42.96
Anthropic: Claude Fable 5 · High
43.70
Anthropic: Claude Fable 5 · Medium
47.55
Spring Prompt benchmark
307
Tournament prediction points
Fixture-by-fixture football predictions frozen before the tournament and graded on outcomes, scorelines, goals, penalties, and cards.
32 fixtures graded from the frozen prediction artifact.
Claude Fable 5
307
source-native benchmark
60
A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
60
source-native benchmark
59.2
Coding Agent Index v1.2 point score
Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.
These are exact agent-plus-model configurations, not bare-model scores.
Claude Code - Fable 5 (max) (with fallback)
59.2
source-native benchmark
62%
Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
62%
source-native benchmark
93%
The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
93%
source-native benchmark
53%
A broad expert-level benchmark of difficult academic reasoning and knowledge questions.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
53%
source-native benchmark
63%
An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
63%
source-native benchmark
95.8%
OpenHands SWE-Bench resolved
Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.
Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.
OpenHands v1.28.0 + claude-fable-5
95.8%
source-native benchmark
85.09%
Direct benchmark score
Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.
Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.
claude-fable-5
85.09%
source-native benchmark
85%
A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
85%
source-native benchmark
99%
A legacy agentic tool-use benchmark for conversational work in a telecom environment.
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
99%
Independent operational reference
Representative configuration: Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback). Lifecycle: active; released 2026-06-09.
Input price
$10.00 / 1M tokens
Output price
$50.00 / 1M tokens
Median output speed
65.0 tok/s
Operational reference only. Price and throughput are not model-quality scores and never affect Spring Prompt comparisons.
Early-user signal
A paraphrased editorial synthesis of 4 Reddit discussions created in the 14 days after launch (9 Jun 2026 to 23 Jun 2026). Anecdotal context only — it never affects scores or rankings.
In the search-indexed launch-window sample, Claude Fable 5 appeared especially capable on difficult code review, architecture, and long-running agent work when tasks were tightly scoped. Reports also repeatedly noted slow interactive responses, high usage cost, and safeguards or product routing that could interrupt the expected model experience, so early results varied substantially by workflow.
Independent early-test signal
An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 9 Jun 2026 up to 23 Jun 2026, and exact model identity was editor reviewed.
Initial field demonstrations showed Fable 5 turning personal sports footage into a tailored visual-analysis overlay and building a navigable, real-scale Yosemite scene from geospatial data. Both artifacts point to strong prompt-to-product work across visual reasoning and procedural 3D, but they are self-selected demonstrations and do not establish routine reliability or failure rates.
These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.
Business Skills V3 · proposed
The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.
Future first-party coverage
These links describe evaluation contracts, not published Claude Fable 5 results.
Model comparisons
Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.
Publication safeguards
Stable family URL
New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.
More from Anthropic