Models / Claude Sonnet 4.6

Anthropic

Claude Sonnet 4.6

29 published results from 4 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Anthropic
Sources
4
Our benchmarks
0
Price
Not yet published

Reported by others

1,478
Overall · rank 6 of 44
Unit
Arena rating, higher is better
Range
1,472 to 1,484
Sample
57813 votes
Configuration
Claude Sonnet 4.6
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,201
Overall · rank 6 of 34
Unit
Arena rating, higher is better
Range
1,196 to 1,206
Sample
134905 votes
Configuration
Claude Sonnet 4.6
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,461
Overall · rank 29 of 177
Unit
Arena rating, higher is better
Range
1,458 to 1,464
Sample
70397 votes
Configuration
Claude Sonnet 4.6
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,482
Business, management and finance · rank 8 of 402
Unit
Arena rating, higher is better
Range
1,476 to 1,488
Sample
14081 votes
Configuration
Claude Sonnet 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,449
Creative writing · rank 20 of 407
Unit
Arena rating, higher is better
Range
1,442 to 1,456
Sample
12132 votes
Configuration
Claude Sonnet 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,506
Expert prompts · rank 13 of 359
Unit
Arena rating, higher is better
Range
1,498 to 1,514
Sample
7372 votes
Configuration
Claude Sonnet 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,475
Instruction following · rank 13 of 409
Unit
Arena rating, higher is better
Range
1,469 to 1,480
Sample
23715 votes
Configuration
Claude Sonnet 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,472
Overall · rank 23 of 409
Unit
Arena rating, higher is better
Range
1,469 to 1,476
Sample
70662 votes
Configuration
Claude Sonnet 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,455
Writing, literature and language · rank 25 of 408
Unit
Arena rating, higher is better
Range
1,449 to 1,461
Sample
17418 votes
Configuration
Claude Sonnet 4.6
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by OpenHands Index
44.5%
Average · rank 23 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.6
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
30.9%
Front end (SWE-Bench Multimodal) · rank 20 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.6
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
50.0%
Greenfield (Commit0) · rank 6 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.6
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
13.3%
Information gathering (GAIA) · rank 32 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.6
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
74.4%
Issue resolution (SWE-Bench) · rank 16 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.6
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
54.0%
Testing (SWT-Bench) · rank 24 of 34
Unit
% resolved, higher is better
Configuration
claude-sonnet-4.6
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by UGI Leaderboard
3.0%
Requested-length error · rank 18 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
2.0%
Requested-length error · rank 7 of 370
Unit
% off the requested word count, lower is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.39
Style adherence · rank 35 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.39
Style adherence · rank 39 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.41
Style adherence · rank 9 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.41
Style adherence · rank 9 of 370
Unit
score from 0 to 1, higher is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
64.5
Writing score · rank 69 of 370
Unit
score out of 100, higher is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
59.3
Writing score · rank 104 of 370
Unit
score out of 100, higher is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
64.2
Writing score · rank 74 of 370
Unit
score out of 100, higher is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
61.2
Writing score · rank 92 of 370
Unit
score out of 100, higher is better
Configuration
claude-sonnet-4.6
Measured
18 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
99.9%
Answer rate · rank 19 of 108
Unit
% of documents, higher is better
Configuration
Claude Sonnet 4.6
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
10.6%
Hallucination rate · rank 65 of 108
Unit
% of summaries, lower is better
Configuration
Claude Sonnet 4.6
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Claude Sonnet 4.6 with