Models / Claude Opus 4.6

Anthropic

Claude Opus 4.6

57 published results from 10 sources. Each card shows where the number comes from and what it does not measure. The overall leaderboard combines them; here each stands alone.

Provider
Anthropic
Sources
10
Our benchmarks
0
Price per million tokens
$5.00 in · $25.00 out
OpenRouter list price, 1 Oct 2026 · 1,000,000-token context

Reported by others

Reported by APEX-Agents
46.3%
Tasks passed · rank 27 of 39
Unit
% of tasks, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
1 Oct 2026
Not shown
Not your firm's documents or tools; graded by rubric, not by a client.
1,494
Overall · rank 4 of 44
Unit
Arena rating, higher is better
Range
1,488 to 1,500
Sample
41260 votes
Configuration
Claude Opus 4.6
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,495
Overall · rank 2 of 44
Unit
Arena rating, higher is better
Range
1,489 to 1,502
Sample
27929 votes
Configuration
Claude Opus 4.6 (high)
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,223
Overall · rank 4 of 34
Unit
Arena rating, higher is better
Range
1,218 to 1,228
Sample
134699 votes
Configuration
Claude Opus 4.6
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,488
Overall · rank 3 of 177
Unit
Arena rating, higher is better
Range
1,486 to 1,491
Sample
80466 votes
Configuration
Claude Opus 4.6
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,494
Overall · rank 2 of 177
Unit
Arena rating, higher is better
Range
1,491 to 1,497
Sample
76191 votes
Configuration
Claude Opus 4.6 (high)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,501
Business, management and finance · rank 6 of 402
Unit
Arena rating, higher is better
Range
1,495 to 1,507
Sample
16124 votes
Configuration
Claude Opus 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,501
Business, management and finance · rank 5 of 402
Unit
Arena rating, higher is better
Range
1,495 to 1,507
Sample
15222 votes
Configuration
Claude Opus 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,478
Creative writing · rank 11 of 407
Unit
Arena rating, higher is better
Range
1,472 to 1,484
Sample
13722 votes
Configuration
Claude Opus 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,501
Creative writing · rank 3 of 407
Unit
Arena rating, higher is better
Range
1,494 to 1,507
Sample
13637 votes
Configuration
Claude Opus 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,533
Expert prompts · rank 9 of 359
Unit
Arena rating, higher is better
Range
1,526 to 1,541
Sample
8315 votes
Configuration
Claude Opus 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,547
Expert prompts · rank 3 of 359
Unit
Arena rating, higher is better
Range
1,539 to 1,555
Sample
6977 votes
Configuration
Claude Opus 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,500
Instruction following · rank 5 of 409
Unit
Arena rating, higher is better
Range
1,495 to 1,505
Sample
27183 votes
Configuration
Claude Opus 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,513
Instruction following · rank 2 of 409
Unit
Arena rating, higher is better
Range
1,508 to 1,519
Sample
24708 votes
Configuration
Claude Opus 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,498
Overall · rank 6 of 409
Unit
Arena rating, higher is better
Range
1,494 to 1,501
Sample
80836 votes
Configuration
Claude Opus 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,505
Overall · rank 2 of 409
Unit
Arena rating, higher is better
Range
1,502 to 1,509
Sample
76518 votes
Configuration
Claude Opus 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,490
Writing, literature and language · rank 7 of 408
Unit
Arena rating, higher is better
Range
1,485 to 1,496
Sample
19764 votes
Configuration
Claude Opus 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,500
Writing, literature and language · rank 3 of 408
Unit
Arena rating, higher is better
Range
1,495 to 1,506
Sample
19461 votes
Configuration
Claude Opus 4.6 (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
31.9
Artificial Analysis Intelligence Index · rank 94 of 314
Unit
index score, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
1 Oct 2026
Not shown
Not business work, and a blend: read the parts for any one task.
26.4
Artificial Analysis Intelligence Index · rank 124 of 314
Unit
index score, higher is better
Configuration
Claude Opus 4.6 (off)
Measured
1 Oct 2026
Not shown
Not business work, and a blend: read the parts for any one task.
89.6%
GPQA Diamond · rank 70 of 276
Unit
% of questions, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
1 Oct 2026
Not shown
Not applied work; multiple-choice science questions.
84.0%
GPQA Diamond · rank 126 of 276
Unit
% of questions, higher is better
Configuration
Claude Opus 4.6 (off)
Measured
1 Oct 2026
Not shown
Not applied work; multiple-choice science questions.
39.9%
Humanity's Last Exam · rank 75 of 314
Unit
% of questions, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
1 Oct 2026
Not shown
Not everyday work; academic questions at the edge of expertise.
19.1%
Humanity's Last Exam · rank 172 of 314
Unit
% of questions, higher is better
Configuration
Claude Opus 4.6 (off)
Measured
1 Oct 2026
Not shown
Not everyday work; academic questions at the edge of expertise.
53.1%
IFBench · rank 106 of 215
Unit
% of instructions, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
1 Oct 2026
Not shown
Not judgement about what an instruction meant.
44.6%
IFBench · rank 142 of 215
Unit
% of instructions, higher is better
Configuration
Claude Opus 4.6 (off)
Measured
1 Oct 2026
Not shown
Not judgement about what an instruction meant.
78.0%
Long-context reasoning (AA-LCR) · rank 98 of 309
Unit
% of questions, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
1 Oct 2026
Not shown
Not retrieval over your own document store.
67.0%
Long-context reasoning (AA-LCR) · rank 186 of 309
Unit
% of questions, higher is better
Configuration
Claude Opus 4.6 (off)
Measured
1 Oct 2026
Not shown
Not retrieval over your own document store.
46.2%
Terminal-Bench Hard · rank 29 of 213
Unit
% of tasks, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
1 Oct 2026
Not shown
Not other harnesses; Artificial Analysis no longer runs it on new models.
48.5%
Terminal-Bench Hard · rank 25 of 213
Unit
% of tasks, higher is better
Configuration
Claude Opus 4.6 (off)
Measured
1 Oct 2026
Not shown
Not other harnesses; Artificial Analysis no longer runs it on new models.
92.1%
Τ²-bench telecom · rank 42 of 214
Unit
% of tasks, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
1 Oct 2026
Not shown
Not your policies or systems; no longer run on new models.
84.8%
Τ²-bench telecom · rank 71 of 214
Unit
% of tasks, higher is better
Configuration
Claude Opus 4.6 (off)
Measured
1 Oct 2026
Not shown
Not your policies or systems; no longer run on new models.
Reported by OpenHands Index
66.7%
Average · rank 4 of 34
Unit
% resolved, higher is better
Configuration
Claude Opus 4.6 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
41.8%
Front end (SWE-Bench Multimodal) · rank 5 of 34
Unit
% resolved, higher is better
Configuration
Claude Opus 4.6 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
56.2%
Greenfield (Commit0) · rank 3 of 34
Unit
% resolved, higher is better
Configuration
Claude Opus 4.6 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
80.0%
Information gathering (GAIA) · rank 7 of 34
Unit
% resolved, higher is better
Configuration
Claude Opus 4.6 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
76.8%
Issue resolution (SWE-Bench) · rank 6 of 34
Unit
% resolved, higher is better
Configuration
Claude Opus 4.6 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
78.8%
Testing (SWT-Bench) · rank 6 of 34
Unit
% resolved, higher is better
Configuration
Claude Opus 4.6 (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by Remote Labor Index
4.2%
Projects done to client standard · rank 7 of 14
Unit
% of projects, higher is better
Configuration
Claude Opus 4.6
Measured
1 Oct 2026
Not shown
Not speed or cost; a small set of projects, so a few points either way are noise.
47.0%
Correct answers · rank 32 of 85
Unit
% of questions, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
27 Aug 2026
Not shown
Not answers grounded in your documents; tests what the model remembers.
Reported by tau2-bench
11.3%
Consistency (pass^4) · rank 15 of 21
Unit
% of tasks, higher is better
Configuration
Claude Opus 4.6 (max reasoning, tau2)
Measured
6 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
27.3%
Task success (pass^1) · rank 14 of 21
Unit
% of tasks, higher is better
Configuration
Claude Opus 4.6 (max reasoning, tau2)
Measured
6 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by UGI Leaderboard
5.0%
Requested-length error · rank 34 of 370
Unit
% off the requested word count, lower is better
Configuration
Claude Opus 4.6 (high)
Measured
6 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
4.0%
Requested-length error · rank 28 of 370
Unit
% off the requested word count, lower is better
Configuration
Claude Opus 4.6 (low reasoning)
Measured
6 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
3.0%
Requested-length error · rank 18 of 370
Unit
% off the requested word count, lower is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
6 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
4.0%
Requested-length error · rank 28 of 370
Unit
% off the requested word count, lower is better
Configuration
Claude Opus 4.6 (medium reasoning)
Measured
6 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.41
Style adherence · rank 7 of 370
Unit
score from 0 to 1, higher is better
Configuration
Claude Opus 4.6 (high)
Measured
6 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.40
Style adherence · rank 30 of 370
Unit
score from 0 to 1, higher is better
Configuration
Claude Opus 4.6 (low reasoning)
Measured
6 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.41
Style adherence · rank 14 of 370
Unit
score from 0 to 1, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
6 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.41
Style adherence · rank 12 of 370
Unit
score from 0 to 1, higher is better
Configuration
Claude Opus 4.6 (medium reasoning)
Measured
6 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
70.9
Writing score · rank 17 of 370
Unit
score out of 100, higher is better
Configuration
Claude Opus 4.6 (high)
Measured
6 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
66.1
Writing score · rank 55 of 370
Unit
score out of 100, higher is better
Configuration
Claude Opus 4.6 (low reasoning)
Measured
6 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
70.7
Writing score · rank 18 of 370
Unit
score out of 100, higher is better
Configuration
Claude Opus 4.6 (max reasoning)
Measured
6 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
69.0
Writing score · rank 30 of 370
Unit
score out of 100, higher is better
Configuration
Claude Opus 4.6 (medium reasoning)
Measured
6 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
99.8%
Answer rate · rank 34 of 108
Unit
% of documents, higher is better
Configuration
Claude Opus 4.6
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
12.2%
Hallucination rate · rank 83 of 108
Unit
% of summaries, lower is better
Configuration
Claude Opus 4.6
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.
Reported by Vending-Bench 2
$8,017.59
Money after a year · rank 9 of 63
Unit
US dollars, higher is better
Configuration
Claude Opus 4.6
Measured
1 Oct 2026
Not shown
Not a real business; one simulated market with set rules.

Compare Claude Opus 4.6 with