Models / Claude Opus 4.7

Anthropic

Claude Opus 4.7

45 published results from 6 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Anthropic
Sources
6
Our benchmarks
0
Price
Not yet published

Reported by others

1,494
Overall · rank 1 of 44
Unit
Arena rating, higher is better
Range
1,488 to 1,501
Sample
22192 votes
Configuration
Claude Opus 4.7
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,493
Overall · rank 1 of 44
Unit
Arena rating, higher is better
Range
1,486 to 1,500
Sample
21957 votes
Configuration
Claude Opus 4.7 (high)
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,212
Overall · rank 5 of 34
Unit
Arena rating, higher is better
Range
1,207 to 1,218
Sample
91394 votes
Configuration
Claude Opus 4.7
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,475
Overall · rank 9 of 177
Unit
Arena rating, higher is better
Range
1,472 to 1,479
Sample
64953 votes
Configuration
Claude Opus 4.7
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,477
Overall · rank 7 of 177
Unit
Arena rating, higher is better
Range
1,474 to 1,480
Sample
63912 votes
Configuration
Claude Opus 4.7 (high)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,495
Business, management and finance · rank 1 of 402
Unit
Arena rating, higher is better
Range
1,489 to 1,502
Sample
13082 votes
Configuration
Claude Opus 4.7
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,508
Business, management and finance · rank 1 of 402
Unit
Arena rating, higher is better
Range
1,501 to 1,514
Sample
12876 votes
Configuration
Claude Opus 4.7 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,484
Creative writing · rank 4 of 407
Unit
Arena rating, higher is better
Range
1,477 to 1,491
Sample
11591 votes
Configuration
Claude Opus 4.7
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,488
Creative writing · rank 1 of 407
Unit
Arena rating, higher is better
Range
1,481 to 1,496
Sample
11381 votes
Configuration
Claude Opus 4.7 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,535
Expert prompts · rank 1 of 359
Unit
Arena rating, higher is better
Range
1,527 to 1,543
Sample
6985 votes
Configuration
Claude Opus 4.7
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,534
Expert prompts · rank 1 of 359
Unit
Arena rating, higher is better
Range
1,526 to 1,542
Sample
6825 votes
Configuration
Claude Opus 4.7 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,494
Instruction following · rank 3 of 409
Unit
Arena rating, higher is better
Range
1,488 to 1,499
Sample
23030 votes
Configuration
Claude Opus 4.7
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,503
Instruction following · rank 2 of 409
Unit
Arena rating, higher is better
Range
1,497 to 1,508
Sample
22448 votes
Configuration
Claude Opus 4.7 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,495
Overall · rank 3 of 409
Unit
Arena rating, higher is better
Range
1,491 to 1,498
Sample
65051 votes
Configuration
Claude Opus 4.7
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,502
Overall · rank 1 of 409
Unit
Arena rating, higher is better
Range
1,498 to 1,506
Sample
64007 votes
Configuration
Claude Opus 4.7 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,485
Writing, literature and language · rank 4 of 408
Unit
Arena rating, higher is better
Range
1,478 to 1,491
Sample
16514 votes
Configuration
Claude Opus 4.7
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,496
Writing, literature and language · rank 1 of 408
Unit
Arena rating, higher is better
Range
1,490 to 1,502
Sample
16137 votes
Configuration
Claude Opus 4.7 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
33.9%
Consistency (pass^5) · rank 2 of 5
Unit
% of tasks, higher is better
Configuration
claude-opus-4.7
Measured
29 May 2026
Not shown
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
51.0%
Customer support task success (pass@1) · rank 2 of 5
Unit
% of tasks, higher is better
Configuration
claude-opus-4.7
Measured
29 May 2026
Not shown
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
54.0%
Shopping assistant task success (pass@1) · rank 1 of 5
Unit
% of tasks, higher is better
Configuration
claude-opus-4.7
Measured
29 May 2026
Not shown
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
53.4%
Task success (pass@1) · rank 2 of 5
Unit
% of tasks, higher is better
Configuration
claude-opus-4.7
Measured
29 May 2026
Not shown
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
55.0%
Travel task success (pass@1) · rank 2 of 5
Unit
% of tasks, higher is better
Configuration
claude-opus-4.7
Measured
29 May 2026
Not shown
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
3.46
User experience · rank 2 of 5
Unit
score from 1 to 5, higher is better
Configuration
claude-opus-4.7
Measured
29 May 2026
Not shown
Not real customer satisfaction: the ratings come from a model-based judge.
Reported by OpenHands Index
69.7%
Average · rank 3 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.7
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
48.5%
Front end (SWE-Bench Multimodal) · rank 3 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.7
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
56.2%
Greenfield (Commit0) · rank 3 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.7
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
81.2%
Information gathering (GAIA) · rank 5 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.7
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
81.6%
Issue resolution (SWE-Bench) · rank 3 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.7
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
80.8%
Testing (SWT-Bench) · rank 5 of 34
Unit
% resolved, higher is better
Configuration
claude-opus-4.7
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by tau2-bench
24.7%
Consistency (pass^4) · rank 7 of 21
Unit
% of tasks, higher is better
Configuration
claude-opus-4.7
Measured
23 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
40.2%
Task success (pass^1) · rank 7 of 21
Unit
% of tasks, higher is better
Configuration
claude-opus-4.7
Measured
23 Jul 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
Claude Opus 4.7 (high)
Measured
16 Apr 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
1.0%
Requested-length error · rank 3 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.7
Measured
16 Apr 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
3.0%
Requested-length error · rank 18 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.7
Measured
16 Apr 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-opus-4.7
Measured
16 Apr 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.41
Style adherence · rank 8 of 370
Unit
score from 0 to 1, higher is better
Configuration
Claude Opus 4.7 (high)
Measured
16 Apr 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.39
Style adherence · rank 43 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.7
Measured
16 Apr 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.40
Style adherence · rank 23 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.7
Measured
16 Apr 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.40
Style adherence · rank 19 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-opus-4.7
Measured
16 Apr 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
56.9
Writing score · rank 119 of 370
Unit
score out of 100, higher is better
Configuration
Claude Opus 4.7 (high)
Measured
16 Apr 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
56.9
Writing score · rank 118 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.7
Measured
16 Apr 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
61.4
Writing score · rank 91 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.7
Measured
16 Apr 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
56.6
Writing score · rank 122 of 370
Unit
score out of 100, higher is better
Configuration
claude-opus-4.7
Measured
16 Apr 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
98.0%
Answer rate · rank 82 of 108
Unit
% of documents, higher is better
Configuration
Claude Opus 4.7
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
12.0%
Hallucination rate · rank 78 of 108
Unit
% of summaries, lower is better
Configuration
Claude Opus 4.7
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Claude Opus 4.7 with