Models / Claude Sonnet 4.5

Anthropic

Claude Sonnet 4.5

54 published results from 6 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Anthropic
Sources
6
Our benchmarks
0
Price
Not yet published

Reported by others

1,449
Overall · rank 20 of 44
Unit
Arena rating, higher is better
Range
1,443 to 1,455
Sample
31455 votes
Configuration
Claude Sonnet 4.5
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,175
Overall · rank 21 of 34
Unit
Arena rating, higher is better
Range
1,170 to 1,180
Sample
127378 votes
Configuration
Claude Sonnet 4.5
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,449
Overall · rank 54 of 177
Unit
Arena rating, higher is better
Range
1,446 to 1,451
Sample
77332 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,447
Overall · rank 54 of 177
Unit
Arena rating, higher is better
Range
1,444 to 1,449
Sample
74579 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,461
Business, management and finance · rank 33 of 402
Unit
Arena rating, higher is better
Range
1,455 to 1,466
Sample
15821 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,459
Business, management and finance · rank 35 of 402
Unit
Arena rating, higher is better
Range
1,453 to 1,464
Sample
16014 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,454
Creative writing · rank 18 of 407
Unit
Arena rating, higher is better
Range
1,448 to 1,460
Sample
12359 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,450
Creative writing · rank 20 of 407
Unit
Arena rating, higher is better
Range
1,444 to 1,456
Sample
12639 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,483
Expert prompts · rank 35 of 359
Unit
Arena rating, higher is better
Range
1,475 to 1,491
Sample
6139 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,498
Expert prompts · rank 17 of 359
Unit
Arena rating, higher is better
Range
1,490 to 1,506
Sample
6061 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,462
Instruction following · rank 28 of 409
Unit
Arena rating, higher is better
Range
1,457 to 1,466
Sample
24284 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,464
Instruction following · rank 24 of 409
Unit
Arena rating, higher is better
Range
1,459 to 1,468
Sample
24600 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,455
Overall · rank 53 of 409
Unit
Arena rating, higher is better
Range
1,452 to 1,458
Sample
82451 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,457
Overall · rank 53 of 409
Unit
Arena rating, higher is better
Range
1,454 to 1,459
Sample
83689 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,454
Writing, literature and language · rank 27 of 408
Unit
Arena rating, higher is better
Range
1,449 to 1,459
Sample
18925 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,452
Writing, literature and language · rank 28 of 408
Unit
Arena rating, higher is better
Range
1,447 to 1,457
Sample
19307 votes
Configuration
Claude Sonnet 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
86.6%
Irrelevance detection · rank 26 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
95.0%
Irrelevance detection · rank 5 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
65.0%
Memory · rank 2 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
5.4%
Memory · rank 89 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
61.4%
Multi-turn tasks · rank 9 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
1.6%
Multi-turn tasks · rank 97 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
73.2%
Overall accuracy · rank 2 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
24.9%
Overall accuracy · rank 89 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
68.8%
Relevance detection · rank 80 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
37.5%
Relevance detection · rank 103 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
88.7%
Single-turn calls (curated) · rank 13 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
59.8%
Single-turn calls (curated) · rank 97 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
81.1%
Single-turn calls (user-contributed) · rank 7 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
46.6%
Single-turn calls (user-contributed) · rank 102 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
81.0%
Web search · rank 6 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
16.0%
Web search · rank 47 of 109
Unit
% correct, higher is better
Configuration
claude-sonnet-4.5
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
Reported by OpenHands Index
53.0%
Average · rank 15 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
36.8%
Front end (SWE-Bench Multimodal) · rank 11 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
12.5%
Greenfield (Commit0) · rank 25 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
72.7%
Information gathering (GAIA) · rank 10 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
74.2%
Issue resolution (SWE-Bench) · rank 17 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
68.8%
Testing (SWT-Bench) · rank 16 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by tau2-bench
48.0%
Consistency (pass^4) · rank 7 of 8
Unit
% of tasks, higher is better
Configuration
claude-sonnet-4.5
Measured
24 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
72.0%
Task success (pass^1) · rank 7 of 8
Unit
% of tasks, higher is better
Configuration
claude-sonnet-4.5
Measured
24 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
10.3%
Consistency (pass^4) · rank 1 of 3
Unit
% of tasks, higher is better
Configuration
claude-sonnet-4.5
Measured
24 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
25.3%
Task success (pass^1) · rank 2 of 3
Unit
% of tasks, higher is better
Configuration
claude-sonnet-4.5
Measured
24 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
39.5%
Consistency (pass^4) · rank 8 of 8
Unit
% of tasks, higher is better
Configuration
claude-sonnet-4.5
Measured
24 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
72.4%
Task success (pass^1) · rank 8 of 8
Unit
% of tasks, higher is better
Configuration
claude-sonnet-4.5
Measured
24 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
64.0%
Consistency (pass^4) · rank 6 of 8
Unit
% of tasks, higher is better
Configuration
claude-sonnet-4.5
Measured
24 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
84.9%
Task success (pass^1) · rank 7 of 8
Unit
% of tasks, higher is better
Configuration
claude-sonnet-4.5
Measured
24 Feb 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by UGI Leaderboard
5.0%
Requested-length error · rank 34 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-sonnet-4.5
Measured
30 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
8.0%
Requested-length error · rank 55 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-sonnet-4.5
Measured
30 Sep 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.38
Style adherence · rank 71 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-sonnet-4.5
Measured
30 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.41
Style adherence · rank 16 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-sonnet-4.5
Measured
30 Sep 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
66.9
Writing score · rank 49 of 370
Unit
score out of 100, higher is better
Configuration
claude-sonnet-4.5
Measured
30 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
64.2
Writing score · rank 73 of 370
Unit
score out of 100, higher is better
Configuration
claude-sonnet-4.5
Measured
30 Sep 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
95.6%
Answer rate · rank 90 of 108
Unit
% of documents, higher is better
Configuration
Claude Sonnet 4.5
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
12.0%
Hallucination rate · rank 78 of 108
Unit
% of summaries, lower is better
Configuration
Claude Sonnet 4.5
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Claude Sonnet 4.5 with