Models / Claude Opus 4.5

Anthropic

Claude Opus 4.5

54 published results from 6 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Anthropic
Sources
6
Our benchmarks
0
Price
Not yet published

Reported by others

1,465
Overall · rank 11 of 44
Unit
Arena rating, higher is better
Range
1,454 to 1,475
Sample
7981 votes
Configuration
Claude Opus 4.5
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,187
Overall · rank 11 of 34
Unit
Arena rating, higher is better
Range
1,181 to 1,193
Sample
61573 votes
Configuration
Claude Opus 4.5
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,459
Overall · rank 35 of 177
Unit
Arena rating, higher is better
Range
1,456 to 1,462
Sample
71633 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,459
Overall · rank 32 of 177
Unit
Arena rating, higher is better
Range
1,456 to 1,463
Sample
36596 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,472
Business, management and finance · rank 15 of 402
Unit
Arena rating, higher is better
Range
1,466 to 1,478
Sample
14086 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,467
Business, management and finance · rank 20 of 402
Unit
Arena rating, higher is better
Range
1,460 to 1,475
Sample
7034 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,461
Creative writing · rank 12 of 407
Unit
Arena rating, higher is better
Range
1,455 to 1,468
Sample
11375 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,469
Creative writing · rank 6 of 407
Unit
Arena rating, higher is better
Range
1,461 to 1,477
Sample
5680 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,503
Expert prompts · rank 13 of 359
Unit
Arena rating, higher is better
Range
1,494 to 1,511
Sample
5682 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,503
Expert prompts · rank 13 of 359
Unit
Arena rating, higher is better
Range
1,491 to 1,515
Sample
2424 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,475
Instruction following · rank 13 of 409
Unit
Arena rating, higher is better
Range
1,470 to 1,480
Sample
21786 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,483
Instruction following · rank 7 of 409
Unit
Arena rating, higher is better
Range
1,477 to 1,490
Sample
9875 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,470
Overall · rank 27 of 409
Unit
Arena rating, higher is better
Range
1,467 to 1,473
Sample
72766 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,474
Overall · rank 21 of 409
Unit
Arena rating, higher is better
Range
1,470 to 1,477
Sample
37468 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,467
Writing, literature and language · rank 15 of 408
Unit
Arena rating, higher is better
Range
1,461 to 1,472
Sample
16876 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,466
Writing, literature and language · rank 14 of 408
Unit
Arena rating, higher is better
Range
1,459 to 1,473
Sample
8564 votes
Configuration
Claude Opus 4.5
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
84.7%
Irrelevance detection · rank 35 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
90.8%
Irrelevance detection · rank 14 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
73.8%
Memory · rank 1 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
1.9%
Memory · rank 102 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
68.4%
Multi-turn tasks · rank 4 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
16.1%
Multi-turn tasks · rank 56 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
77.5%
Overall accuracy · rank 1 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
33.5%
Overall accuracy · rank 57 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
62.5%
Relevance detection · rank 91 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
68.8%
Relevance detection · rank 80 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
88.6%
Single-turn calls (curated) · rank 15 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
89.7%
Single-turn calls (curated) · rank 5 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
79.8%
Single-turn calls (user-contributed) · rank 14 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
76.0%
Single-turn calls (user-contributed) · rank 35 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
84.5%
Web search · rank 1 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
13.0%
Web search · rank 52 of 109
Unit
% correct, higher is better
Configuration
claude-opus-4.5
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
Reported by OpenHands Index
60.6%
Average · rank 8 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
41.2%
Front end (SWE-Bench Multimodal) · rank 6 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
37.5%
Greenfield (Commit0) · rank 10 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
69.1%
Information gathering (GAIA) · rank 13 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
76.6%
Issue resolution (SWE-Bench) · rank 8 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
78.5%
Testing (SWT-Bench) · rank 7 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.5
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by tau2-bench
70.0%
Consistency (pass^4) · rank 2 of 8
Unit
% of tasks, higher is better
Configuration
claude-opus-4.5
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
84.0%
Task success (pass^1) · rank 1 of 8
Unit
% of tasks, higher is better
Configuration
claude-opus-4.5
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
11.3%
Consistency (pass^4) · rank 15 of 21
Unit
% of tasks, higher is better
Configuration
claude-opus-4.5
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
24.7%
Task success (pass^1) · rank 17 of 21
Unit
% of tasks, higher is better
Configuration
claude-opus-4.5
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
51.8%
Consistency (pass^4) · rank 2 of 8
Unit
% of tasks, higher is better
Configuration
claude-opus-4.5
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
79.6%
Task success (pass^1) · rank 3 of 8
Unit
% of tasks, higher is better
Configuration
claude-opus-4.5
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
78.1%
Consistency (pass^4) · rank 2 of 8
Unit
% of tasks, higher is better
Configuration
claude-opus-4.5
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
92.3%
Task success (pass^1) · rank 2 of 8
Unit
% of tasks, higher is better
Configuration
claude-opus-4.5
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by UGI Leaderboard
3.0%
Requested-length error · rank 18 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.5
Measured
24 Nov 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
4.0%
Requested-length error · rank 28 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.5
Measured
24 Nov 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.37
Style adherence · rank 108 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.5
Measured
24 Nov 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.39
Style adherence · rank 40 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.5
Measured
24 Nov 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
70.4
Writing score · rank 19 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.5
Measured
24 Nov 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
70.3
Writing score · rank 20 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.5
Measured
24 Nov 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
98.7%
Answer rate · rank 74 of 108
Unit
% of documents, higher is better
Configuration
Claude Opus 4.5
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
10.9%
Hallucination rate · rank 70 of 108
Unit
% of summaries, lower is better
Configuration
Claude Opus 4.5
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Claude Opus 4.5 with