Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of expert-written exam questions answered correctly, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
405 configurations of 242 models
Unit
% of questions
Licence
Artificial Analysis commercial data licence

Humanity's Last Exam: Gemma 3 270M

Top 15 of 242 results · % of questions, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Claude Opus 5.5 (max reasoning)Anthropic 61.4%
  2. 2 Claude Fable 5.1 (max reasoning)Anthropic 59.1%
  3. 3 Gemini 4 Argon (high reasoning)Google 57.1%
  4. 4 Claude Fable 5 (max reasoning)Anthropic 55.5%
  5. 5 Claude Sonnet 5.5 (max reasoning)Anthropic 55.0%
  6. 6 Claude Opus 5 (max reasoning)Anthropic 54.9%
  7. 7 GPT-6 Astra (max reasoning)OpenAI 54.7%
  8. 8 GPT-6.1 Sol (max reasoning)OpenAI 52.9%
  9. 9 GPT-5.6 Sol (max reasoning)OpenAI 49.5%
  10. 10 MiMo-V2.6-ProXiaomi 49.4%
  11. 11 Claude Opus 4.8 (max reasoning)Anthropic 48.7%
  12. 11 Muse Spark 1.3 (max reasoning)Meta 48.7%
  13. 13 Gemini 3.7 Flash (high reasoning)Google 47.9%
  14. 13 GPT-6 Sol (max reasoning)OpenAI 47.9%
  15. 15 Gemini 3.8 Flash (high reasoning)Google 47.8%
  16. 229 Gemma 3 270MGoogle 3.6%

Full results

Artificial Analysis: Humanity's Last Exam, % of questions, higher is better
#ModelHumanity's Last Exam
% of questions, higher is better
Price
$ per million tokens, in / out
1 Claude Opus 5.5 (max reasoning)Anthropic · best of 5 settings
61.4%
$4 / $20
2 Claude Fable 5.1 (max reasoning)Anthropic · best of 5 settings
59.1%
$10 / $50
3 Gemini 4 Argon (high reasoning)Google
57.1%
–
4 Claude Fable 5 (max reasoning)Anthropic
55.5%
$10 / $50
5 Claude Sonnet 5.5 (max reasoning)Anthropic · best of 5 settings
55.0%
$2 / $10
6 Claude Opus 5 (max reasoning)Anthropic · best of 5 settings
54.9%
$5 / $25
7 GPT-6 Astra (max reasoning)OpenAI · best of 5 settings
54.7%
$10 / $50
8 GPT-6.1 Sol (max reasoning)OpenAI · best of 5 settings
52.9%
$2 / $10
9 GPT-5.6 Sol (max reasoning)OpenAI · best of 6 settings
49.5%
$4 / $20
10 MiMo-V2.6-ProXiaomi
49.4%
$0.43 / $0.87
11 Claude Opus 4.8 (max reasoning)Anthropic
48.7%
$5 / $25
11 Muse Spark 1.3 (max reasoning)Meta · best of 2 settings
48.7%
$1.25 / $4.25
13 Gemini 3.7 Flash (high reasoning)Google · best of 3 settings
47.9%
$1.50 / $7.50
13 GPT-6 Sol (max reasoning)OpenAI · best of 6 settings
47.9%
$2 / $10
15 Gemini 3.8 Flash (high reasoning)Google · best of 3 settings
47.8%
$1.50 / $7.50
16 Gemini 3.1 Pro PreviewGoogle
47.0%
$2 / $12
17 Kimi K3 (max reasoning)Moonshot AI · best of 2 settings
46.9%
$3 / $15
18 Muse Spark 1.1 (extra-high reasoning)Meta
46.2%
$1.25 / $4.25
19 GPT-5.5 (extra-high reasoning)OpenAI · best of 5 settings
45.8%
$5 / $30
20 Muse Spark 1.2 (extra-high reasoning)Meta
45.5%
$1.25 / $4.25
21 Claude Haiku 5.5 (max reasoning)Anthropic · best of 5 settings
44.4%
$0.10 / $0.50
22 Grok 4.6 (extra-high reasoning)xAI · best of 4 settings
44.1%
$2 / $6
23 GPT-5.4 (extra-high reasoning)OpenAI · best of 3 settings
43.7%
$2.50 / $15
24 Qwen3.8-Max (0902)Alibaba
43.1%
$2 / $6
24 Grok 4.7 (extra-high reasoning)xAI · best of 3 settings
43.1%
$2 / $6
26 Qwen3.8-Max (0803)Alibaba
43.0%
–
27 GPT-5.6 Terra (max reasoning)OpenAI · best of 6 settings
42.9%
$2 / $12
28 Gemini 3.5 Flash (high reasoning)Google · best of 3 settings
42.7%
$1.50 / $9
28 Grok 4.5 (high reasoning)xAI
42.7%
$2 / $6
30 GPT-5.3-Codex (extra-high reasoning)OpenAI
42.5%
$1.75 / $14
31 Qwen3.8-2.4T-A95BAlibaba
42.4%
$2 / $6
32 Claude Opus 4.7 (max reasoning)Anthropic · best of 2 settings
42.3%
$5 / $25
32 GLM-5.3 (max reasoning)Z.ai · best of 2 settings
42.3%
$1.40 / $4.40
34 Claude Sonnet 5 (max reasoning)Anthropic · best of 6 settings
41.3%
$2 / $10
35 GLM-5.2 (max reasoning)Z.ai · best of 2 settings
41.1%
$1.40 / $4.40
36 DeepSeek-V4-Pro (0813, max reasoning)DeepSeek · best of 2 settings
41.0%
$1.32 / $3.96
37 Gemini 3.6 Flash (high reasoning)Google
40.8%
$1.50 / $7.50
38 Muse SparkMeta
40.7%
–
39 Qwen3.7-MaxAlibaba
40.5%
$1.48 / $4.42
40 Claude Opus 4.6 (max reasoning)Anthropic · best of 2 settings
39.9%
$5 / $25
40 GLM-5.3-FlashZ.ai
39.9%
$0.15 / $0.50
42 Gemini 3 Pro Preview (high reasoning)Google · best of 2 settings
39.7%
–
43 GPT-5.6 Luna (max reasoning)OpenAI · best of 6 settings
39.5%
$0.20 / $1.20
44 DeepSeek-V4.1-Flash (max reasoning)DeepSeek · best of 2 settings
39.2%
$0.30 / $1.20
45 MiniMax-M3MiniMax
39.0%
$0.30 / $1.20
46 DeepSeek-V4-Flash (0731, max reasoning)DeepSeek
38.6%
$0.14 / $0.28
47 GPT-6 Luna (max reasoning)OpenAI · best of 6 settings
38.5%
$0.10 / $0.50
48 Grok Build 0.1xAI
38.3%
$1 / $2
49 Qwen3.8-Flash-NextAlibaba
38.0%
–
50 GPT-5.2 (extra-high reasoning)OpenAI · best of 3 settings
37.7%
$1.75 / $14
51 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek · best of 3 settings
37.5%
$1.42 / $2.83
51 Kimi K2.6Moonshot AI · best of 2 settings
37.5%
$0.95 / $4
53 Grok 4.3 (high reasoning)xAI · best of 4 settings
37.2%
$1.25 / $2.50
54 Gemini 3 Flash Preview (reasoning on)Google · best of 2 settings
36.6%
$0.50 / $3
55 GPT-5.2-Codex (extra-high reasoning)OpenAI
35.7%
$1.75 / $14
55 MiMo-V2.5-Pro (reasoning on)Xiaomi · best of 2 settings
35.7%
$0.43 / $0.87
57 Qwen3.7-PlusAlibaba
35.6%
$0.32 / $1.28
58 MiMo-V2.6-FlashXiaomi
35.1%
$0.14 / $0.28
59 Mistral Large 4Mistral AI
35.0%
$1.36 / $4.18
59 Kimi K2.7 CodeMoonshot AI
35.0%
$0.95 / $4
61 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek · best of 3 settings
34.8%
$0.14 / $0.28
62 DeepSeek-V4-Flash-Vision-Exp (max reasoning)DeepSeek
34.5%
$0.44 / $1.32
62 Grok 4.20xAI · best of 2 settings
34.5%
$1.25 / $2.50
64 Qwen3.8-27B (extra-high reasoning)Alibaba · best of 4 settings
33.9%
$0.50 / $3
65 Claude Sonnet 4.6 (max reasoning)Anthropic · best of 2 settings
33.6%
$3 / $15
66 Hy3Tencent
33.5%
$0.14 / $0.58
67 Inkling SmallThinking Machines
33.3%
$0.45 / $1.20
68 Grok 4.20 (0309, reasoning)xAI
32.4%
–
69 Inkling (extra-high reasoning)Thinking Machines
31.9%
$0.95 / $4.05
70 Qwen3.6-Max-PreviewAlibaba
30.8%
$1.03 / $6.16
71 Kimi K2.5Moonshot AI · best of 2 settings
30.7%
$0.57 / $2.85
72 Claude Opus 4.5 (reasoning on)Anthropic · best of 2 settings
30.1%
$5 / $25
72 GLM-5.1Z.ai · best of 2 settings
30.1%
$1.38 / $4.40
74 MiniMax-M2.7MiniMax
29.6%
$0.30 / $1.20
75 GLM-5Z.ai · best of 2 settings
29.3%
$0.95 / $2.55
76 Solar Pro 4Upstage
29.2%
$0.09 / $0.36
77 Qwen3.5-397B-A17BAlibaba · best of 2 settings
29.0%
$0.55 / $3.50
78 DeepSeek-V3.2-SpecialeDeepSeek
28.7%
–
79 GPT-5.1 (high reasoning)OpenAI · best of 2 settings
28.5%
$1.25 / $10
79 GPT-5 (high reasoning)OpenAI · best of 4 settings
28.5%
$1.25 / $10
81 GPT-5.4 nano (extra-high reasoning)OpenAI · best of 3 settings
28.3%
$0.20 / $1.25
82 GPT-5.4 mini (extra-high reasoning)OpenAI · best of 3 settings
28.1%
$0.75 / $4.50
83 Qwen3-Max-ThinkingAlibaba
28.0%
$0.78 / $3.90
84 Qwen3.6-PlusAlibaba
27.8%
$0.33 / $1.95
84 GPT-5-Codex (high reasoning)OpenAI
27.8%
–
84 Hy3 Preview (reasoning on)Tencent · best of 2 settings
27.8%
$0.18 / $0.60
84 GLM-5-TurboZ.ai
27.8%
$1.20 / $4
88 GLM-4.7Z.ai · best of 2 settings
27.4%
$0.54 / $1.98
89 MiMo-V2.5Xiaomi
27.2%
$0.17 / $0.34
90 Grok 4xAI
26.7%
–
91 GPT-5.1-Codex (high reasoning)OpenAI
25.7%
$1.25 / $10
92 Qwen3.5-122B-A10BAlibaba · best of 2 settings
25.2%
$0.26 / $2.08
93 DeepSeek-V3.2 (reasoning on)DeepSeek · best of 2 settings
24.6%
$0.30 / $0.96
94 Grok 4.20 (0309, non-reasoning, no reasoning)xAI
24.5%
–
95 Qwen3.5-27BAlibaba · best of 2 settings
23.9%
$0.27 / $2.16
96 Kimi K2 ThinkingMoonshot AI
23.8%
$0.60 / $2.50
97 Gemma 4 31BGoogle · best of 2 settings
23.6%
$0.14 / $0.40
98 MiniMax-M2.1MiniMax
23.2%
$0.30 / $1.20
99 Qwen3.6-27BAlibaba · best of 2 settings
23.1%
$0.30 / $3.20
100 Gemini 2.5 ProGoogle
22.5%
$1.25 / $10
101 Qwen3.6-35B-A3BAlibaba · best of 2 settings
22.2%
$0.10 / $1
102 Muse Glimmer (high reasoning)Meta
22.0%
–
103 GPT-5.5 Instant (2026-05-26)OpenAI · best of 2 settings
21.6%
–
104 GPT-5 mini (high reasoning)OpenAI · best of 3 settings
21.5%
$0.25 / $2
105 Step 3.5 FlashStepFun
21.1%
$0.10 / $0.30
106 Qwen3.5-35B-A3BAlibaba · best of 2 settings
21.0%
$0.16 / $1.30
107 Nemotron 3 Super 120B A12BNVIDIA
20.8%
$0.085 / $0.40
108 MiniMax-M2.5MiniMax
20.5%
$0.30 / $1.20
109 o3OpenAI
20.1%
$2 / $8
110 gpt-oss-120b (high reasoning)OpenAI · best of 2 settings
19.6%
$0.15 / $0.60
111 Gemma 4 26B A4BGoogle · best of 2 settings
19.3%
$0.10 / $0.30
111 Grok 4.1 Fast (reasoning)xAI
19.3%
–
113 Grok 4 Fast (reasoning)xAI
19.1%
–
114 Gemini 3.5 Flash-LiteGoogle
18.8%
$0.30 / $2.50
115 GPT-5.1-Codex-Mini (high reasoning)OpenAI
18.5%
$0.25 / $2
116 Claude Sonnet 4.5 (reasoning on)Anthropic · best of 2 settings
17.8%
$3 / $15
117 Gemini 3.1 Flash-Lite PreviewGoogle
17.2%
$0.25 / $1.50
118 Mercury 2Inception
17.1%
$0.25 / $0.75
118 GLM-5V-TurboZ.ai
17.1%
$1.20 / $4
120 o4-mini (high reasoning)OpenAI
16.5%
$1.10 / $4.40
121 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek · best of 2 settings
16.4%
$0.27 / $1
122 Qwen3-235B-A22B-Thinking-2507Alibaba
15.9%
$0.30 / $3
123 Trinity Large ThinkingArcee AI
15.8%
$0.25 / $0.80
123 DeepSeek-R1-0528DeepSeek
15.8%
$0.50 / $2.18
125 Gemma 4 12BGoogle · best of 2 settings
15.7%
–
126 Qwen3.5-9BAlibaba · best of 2 settings
14.9%
$0.10 / $0.15
126 Qwen3.5-Omni-PlusAlibaba
14.9%
–
126 DeepSeek-V3.2-Exp (reasoning on)DeepSeek · best of 2 settings
14.9%
$0.27 / $0.41
129 GLM-4.6 (reasoning on)Z.ai · best of 2 settings
14.5%
$0.50 / $2
130 DeepSeek-V3.1 (reasoning on)DeepSeek · best of 2 settings
14.3%
$0.55 / $1.65
131 Gemini 2.5 Flash Preview (09-2025, reasoning on)Google · best of 2 settings
13.8%
–
131 Mistral Medium 3.5Mistral AI
13.8%
$1.50 / $7.50
133 MiniMax-M2MiniMax
13.7%
$0.30 / $1.20
134 GLM-4.5Z.ai
13.0%
$0.60 / $2.20
135 Qwen3-Max-Thinking-PreviewAlibaba
12.7%
–
136 Qwen3-Next-80B-A3B-ThinkingAlibaba
12.6%
$0.15 / $1.20
137 Claude Opus 4.1 (reasoning on)Anthropic
12.5%
$15 / $75
138 Gemini 2.5 Flash (reasoning on)Google · best of 2 settings
12.1%
$0.30 / $2.50
139 o3-mini (high reasoning)OpenAI · best of 2 settings
12.0%
$1.10 / $4.40
140 Qwen3-MaxAlibaba
11.9%
$0.78 / $3.90
140 Qwen3-VL-235B-A22B-ThinkingAlibaba
11.9%
$0.40 / $4
142 Nova 2 Lite (high reasoning)Amazon · best of 4 settings
11.6%
$0.30 / $2.50
143 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA · best of 2 settings
11.4%
$0.05 / $0.20
144 Qwen3-235B-A22B-Instruct-2507Alibaba
11.1%
$0.15 / $0.75
145 Qwen3-235B-A22B (reasoning on)Alibaba · best of 2 settings
11.0%
$0.46 / $1.82
145 gpt-oss-20b (high reasoning)OpenAI · best of 2 settings
11.0%
$0.03 / $0.15
147 DiffusionGemma 26B A4BGoogle
10.8%
–
148 Claude Sonnet 4 (reasoning on)Anthropic · best of 2 settings
10.7%
$3 / $15
149 Claude Haiku 4.5 (reasoning on)Anthropic · best of 2 settings
10.4%
$1 / $5
150 Qwen3-30B-A3B-Thinking-2507Alibaba
10.3%
$0.20 / $2.40
150 Magistral Medium 1.2Mistral AI
10.3%
–
150 Solar Pro 3Upstage
10.3%
$0.15 / $0.60
153 Qwen3-Coder-NextAlibaba
10.1%
$0.18 / $0.90
153 Qwen3-Max-PreviewAlibaba
10.1%
–
153 Qwen3-VL-32B-ThinkingAlibaba
10.1%
–
156 Qwen3.5-4B (reasoning on)Alibaba · best of 2 settings
9.9%
–
156 Mistral Small 4Mistral AI · best of 2 settings
9.9%
$0.15 / $0.60
158 Granite 4.2 8BIBM
9.7%
$0.06 / $0.25
159 GLM-4.6V (reasoning on)Z.ai · best of 2 settings
9.6%
$0.30 / $0.90
160 GPT-5 nano (high reasoning)OpenAI · best of 3 settings
9.5%
$0.05 / $0.40
161 Qwen3-VL-30B-A3B-ThinkingAlibaba
8.9%
$0.29 / $1
162 DeepSeek-R1DeepSeek
8.5%
$0.70 / $2.50
163 Grok Code Fast 1xAI
8.0%
–
164 Qwen3.5-Omni-FlashAlibaba
7.6%
–
164 Qwen3-Next-80B-A3B-InstructAlibaba
7.6%
$0.10 / $1.10
164 GLM-4.7-FlashZ.ai · best of 2 settings
7.6%
$0.06 / $0.40
167 Qwen3-Omni-30B-A3B-ThinkingAlibaba
7.5%
–
168 Qwen3-32B (reasoning on)Alibaba · best of 2 settings
7.4%
$0.14 / $0.40
168 Kimi K2 (0711)Moonshot AI
7.4%
$0.57 / $2.30
170 Gemini 2.5 Flash-Lite Preview (09-2025, reasoning on)Google · best of 2 settings
7.0%
–
170 o1OpenAI
7.0%
$15 / $60
170 GLM-4.5-AirZ.ai
7.0%
$0.14 / $0.86
173 Qwen3-30B-A3B-Instruct-2507Alibaba
6.9%
$0.09 / $0.30
174 Qwen3-VL-32B-InstructAlibaba
6.8%
$0.10 / $0.42
174 Gemini 2.5 Flash-Lite (reasoning on)Google · best of 2 settings
6.8%
$0.10 / $0.40
176 Qwen3-VL-235B-A22B-InstructAlibaba
6.6%
$0.30 / $1.50
176 GPT-5 ChatOpenAI
6.6%
–
178 Magistral Small 1.2Mistral AI
6.4%
–
178 Kimi K2 (0905)Moonshot AI
6.4%
$0.60 / $2.50
180 Qwen3-VL-30B-A3B-InstructAlibaba
6.3%
$0.15 / $0.60
180 GLM-4.5V (reasoning on)Z.ai · best of 2 settings
6.3%
$0.60 / $1.80
182 Qwen3-30B-A3B (reasoning on)Alibaba · best of 2 settings
6.2%
$0.12 / $0.50
182 Qwen3-4B-Thinking-2507Alibaba
6.2%
–
184 Llama 3.2 1B InstructMeta
5.5%
$0.027 / $0.20
185 Ministral 3 3BMistral AI
5.4%
$0.10 / $0.10
186 Gemma 3 4BGoogle
5.3%
$0.05 / $0.10
186 Llama 3.1 8B InstructMeta
5.3%
$0.05 / $0.08
186 Llama 3.2 3B InstructMeta
5.3%
$0.05 / $0.33
189 Qwen3.5-0.8B (no reasoning)Alibaba · best of 2 settings
5.1%
–
189 Grok 4.1 Fast (non-reasoning, no reasoning)xAI
5.1%
–
191 Qwen3.5-2B (no reasoning)Alibaba · best of 2 settings
5.0%
–
191 GPT-4.1 miniOpenAI
5.0%
$0.40 / $1.60
193 Llama 4 MaverickMeta
4.9%
$0.27 / $0.85
194 Gemma 4 E2BGoogle · best of 2 settings
4.8%
–
194 Gemma 4 E4B (no reasoning)Google · best of 2 settings
4.8%
–
196 DeepSeek-V3-0324DeepSeek
4.7%
$0.25 / $1
196 Mistral Medium 3.1Mistral AI
4.7%
$0.40 / $2
198 Qwen3-Omni-30B-A3B-InstructAlibaba
4.6%
–
198 Qwen3-VL-4B-ThinkingAlibaba
4.6%
–
198 Nova Micro 1.0Amazon
4.6%
$0.035 / $0.14
198 Ministral 3 14BMistral AI
4.6%
$0.20 / $0.20
202 Qwen3-14B (reasoning on)Alibaba · best of 2 settings
4.5%
$0.12 / $0.24
202 Qwen3-4B-Instruct-2507Alibaba
4.5%
–
202 Qwen3-Coder-480B-A35BAlibaba
4.5%
$0.35 / $1.50
202 Llama 3.1 70B InstructMeta
4.5%
$0.40 / $0.40
202 Grok 4 Fast (non-reasoning, no reasoning)xAI
4.5%
–
207 Gemma 3 27BGoogle
4.4%
$0.12 / $0.20
207 Reka Flash 3Reka AI
4.4%
$0.10 / $0.20
209 Nova Lite 1.0Amazon
4.3%
$0.06 / $0.24
209 Ministral 3 8BMistral AI
4.3%
$0.15 / $0.15
209 Mistral Small 3.1 24BMistral AI
4.3%
$0.35 / $0.56
209 Mistral Small 3.2 24BMistral AI
4.3%
$0.094 / $0.25
213 Gemma 3 12BGoogle
4.2%
$0.05 / $0.15
213 Mistral Large 3Mistral AI
4.2%
$0.50 / $1.50
213 GPT-4.1OpenAI
4.2%
$2 / $8
213 GPT-4o mini (2024-07-18)OpenAI
4.2%
$0.15 / $0.60
217 Mistral Medium 3Mistral AI
4.1%
$0.40 / $2
218 Command ACohere
4.0%
$2.50 / $10
218 Mixtral 8x22B InstructMistral AI
4.0%
$2 / $6
220 Qwen3-8B (reasoning on)Alibaba · best of 2 settings
3.9%
$0.12 / $0.46
221 Qwen3-Coder-30B-A3B-InstructAlibaba
3.8%
$0.07 / $0.28
221 Qwen3-VL-8B-ThinkingAlibaba
3.8%
$0.18 / $2.10
221 Llama 4 ScoutMeta
3.8%
$0.18 / $0.59
221 Phi-4Microsoft
3.8%
$0.07 / $0.14
221 Devstral MediumMistral AI
3.8%
–
221 Devstral Small 1.1Mistral AI
3.8%
–
221 Mistral Small 3Mistral AI
3.8%
$0.05 / $0.08
221 GPT-4.1 nanoOpenAI
3.8%
$0.10 / $0.40
229 Qwen2.5-72B-InstructAlibaba
3.6%
$0.36 / $0.40
229 Qwen3-VL-4B-InstructAlibaba
3.6%
–
229 Gemma 3 270MGoogle
3.6%
–
229 Llama 3.3 70B InstructMeta
3.6%
$0.59 / $0.79
229 Devstral 2Mistral AI
3.6%
$0.40 / $2
234 Qwen2.5-Coder-32B-InstructAlibaba
3.5%
$0.66 / $1
234 Devstral Small 2Mistral AI
3.5%
–
236 Nova Pro 1.0Amazon
3.2%
$0.80 / $3.20
237 DeepSeek-V3DeepSeek
2.9%
$0.26 / $1.03
237 Mistral Large 2 (2407)Mistral AI
2.9%
$2 / $6
239 Qwen3-VL-8B-InstructAlibaba
2.7%
$0.12 / $0.46
240 Kimi Linear 48B A3B InstructMoonshot AI
2.5%
–
241 GPT-4o (2024-08-06)OpenAI
2.3%
$2.50 / $10
242 GPT-4o (2024-05-13)OpenAI
1.8%
$5 / $15

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of expert-written exam questions answered correctly, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not everyday work; academic questions at the edge of expertise.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.