Benchmarks / SimpleQA Verified (Epoch AI)

Reported by SimpleQA Verified (Epoch AI)

SimpleQA Verified (Epoch AI)

Share of short factual questions answered correctly without search, as run by Epoch AI.

Last updated 29 Sep 2026

Results dated
10 Aug 2026 to 29 Sep 2026
Results
85 configurations of 76 models
Unit
% of questions
Licence
Creative Commons Attribution 4.0 International

Correct answers: Gemini 3.6 Flash

Top 15 of 76 results · % of questions, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 GPT-6 Astra (max reasoning)OpenAI 75.6%
  2. 2 GPT-6.1 Sol (max reasoning)OpenAI 73.9%
  3. 3 Gemini 3.1 Pro Preview (high reasoning)Google 73.5%
  4. 4 Claude Opus 5.5 (max reasoning)Anthropic 72.2%
  5. 5 Claude Fable 5.1 (max reasoning)Anthropic 70.8%
  6. 6 Claude Fable 5 (extra-high reasoning)Anthropic 70.7%
  7. 7 Gemini 3.8 Flash (high reasoning)Google 69.7%
  8. 7 GPT-5.6 Sol (max reasoning)OpenAI 69.7%
  9. 9 Gemini 3.7 Flash (high reasoning)Google 69.2%
  10. 10 Gemini 3 Flash Preview (high reasoning)Google 66.8%
  11. 11 Gemini 3.5 Flash (high reasoning)Google 66.2%
  12. 11 Gemini 3.6 Flash (high reasoning)Google 66.2%
  13. 13 GPT-5.5 (extra-high reasoning)OpenAI 63.0%
  14. 14 GPT-6 Sol (max reasoning)OpenAI 60.7%
  15. 15 Muse Spark 1.2 (extra-high reasoning)Meta 60.3%

Full results

SimpleQA Verified: correct answers, % of questions, higher is better
#ModelCorrect answers
% of questions, higher is better
Price
$ per million tokens, in / out
1 GPT-6 Astra (max reasoning)OpenAI
75.6%
$10 / $50
2 GPT-6.1 Sol (max reasoning)OpenAI
73.9%
$2 / $10
3 Gemini 3.1 Pro Preview (high reasoning)Google
73.5%
$2 / $12
4 Claude Opus 5.5 (max reasoning)Anthropic
72.2%
$4 / $20
5 Claude Fable 5.1 (max reasoning)Anthropic
70.8%
$10 / $50
6 Claude Fable 5 (extra-high reasoning)Anthropic
70.7%
$10 / $50
7 Gemini 3.8 Flash (high reasoning)Google
69.7%
$1.50 / $7.50
7 GPT-5.6 Sol (max reasoning)OpenAI
69.7%
$4 / $20
9 Gemini 3.7 Flash (high reasoning)Google
69.2%
$1.50 / $7.50
10 Gemini 3 Flash Preview (high reasoning)Google
66.8%
$0.50 / $3
11 Gemini 3.5 Flash (high reasoning)Google
66.2%
$1.50 / $9
11 Gemini 3.6 Flash (high reasoning)Google
66.2%
$1.50 / $7.50
13 GPT-5.5 (extra-high reasoning)OpenAI
63.0%
$5 / $30
14 GPT-6 Sol (max reasoning)OpenAI
60.7%
$2 / $10
15 Muse Spark 1.2 (extra-high reasoning)Meta
60.3%
$1.25 / $4.25
16 Claude Opus 5 (max reasoning)Anthropic
59.9%
$5 / $25
17 Muse Spark 1.1Meta
57.8%
$1.25 / $4.25
18 Qwen3.7-MaxAlibaba
55.8%
$1.48 / $4.42
19 Claude Opus 4.8 (max reasoning)Anthropic
53.0%
$5 / $25
20 DeepSeek-V4-Pro (0813, max reasoning)DeepSeek
52.9%
$1.32 / $3.96
21 Qwen3.6-Max-PreviewAlibaba
52.0%
$1.03 / $6.16
22 Claude Opus 4.7 (extra-high reasoning)Anthropic
51.7%
$5 / $25
23 Kimi K3 (max reasoning)Moonshot AI
50.6%
$3 / $15
24 GPT-5 (high reasoning)OpenAI
50.1%
$1.25 / $10
25 o3 (high reasoning)OpenAI
49.4%
$2 / $8
26 Grok 4.6 (high reasoning)xAI · best of 2 settings
49.3%
$2 / $6
27 Qwen3-MaxAlibaba
48.7%
$0.78 / $3.90
28 Grok 4.5 (high reasoning)xAI
48.3%
$2 / $6
29 GPT-5.1 (high reasoning)OpenAI
48.0%
$1.25 / $10
30 Qwen3.8-Max (0902, extra-high reasoning)Alibaba
47.3%
$2 / $6
31 Claude Opus 4.6 (max reasoning)Anthropic
47.0%
$5 / $25
31 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek
47.0%
$1.42 / $2.83
33 Claude Sonnet 5.5 (max reasoning)Anthropic
46.5%
$2 / $10
34 GPT-5.4 Pro (extra-high reasoning)OpenAI
46.3%
$30 / $180
35 Qwen3.8-Max (0803, extra-high reasoning)Alibaba
45.8%
–
36 Claude Opus 4.5 (32k reasoning budget)Anthropic
45.7%
$5 / $25
37 GPT-5.4 (extra-high reasoning)OpenAI
45.1%
$2.50 / $15
38 Qwen3.6-PlusAlibaba
44.1%
$0.33 / $1.95
39 GPT-5.6 Terra (max reasoning)OpenAI
43.2%
$2 / $12
40 GPT-6 Luna (max reasoning)OpenAI
41.4%
$0.10 / $0.50
41 o1 (high reasoning)OpenAI
41.1%
$15 / $60
42 GPT-5.6 Luna (max reasoning)OpenAI
41.0%
$0.20 / $1.20
42 GLM-5.3 (max reasoning)Z.ai
41.0%
$1.40 / $4.40
44 Qwen3-235B-A22B-Thinking-2507Alibaba
40.4%
$0.30 / $3
45 Inkling (extra-high reasoning)Thinking Machines
40.3%
$0.95 / $4.05
46 GPT-5.2 (extra-high reasoning)OpenAI · best of 4 settings
37.1%
$1.75 / $14
47 Kimi K2.7 CodeMoonshot AI
36.5%
$0.95 / $4
48 Claude Sonnet 4.6 (high reasoning)Anthropic · best of 2 settings
35.5%
$3 / $15
49 Kimi K2.6Moonshot AI
34.9%
$0.95 / $4
50 Kimi K2.5Moonshot AI
34.3%
$0.57 / $2.85
51 GLM-5.2 (max reasoning)Z.ai
34.2%
$1.40 / $4.40
52 GLM-5.1Z.ai
34.0%
$1.38 / $4.40
53 Claude Sonnet 5 (max reasoning)Anthropic · best of 2 settings
33.7%
$2 / $10
54 DeepSeek-V4-Flash (0731, max reasoning)DeepSeek
33.6%
$0.14 / $0.28
55 Grok 4.3 (high reasoning)xAI
33.2%
$1.25 / $2.50
56 GLM-4.7Z.ai
32.2%
$0.54 / $1.98
57 GPT-4.1OpenAI
31.1%
$2 / $8
58 Claude Sonnet 4.5 (59k reasoning budget)Anthropic · best of 2 settings
30.7%
$3 / $15
59 Grok 4.20 (0309, reasoning)xAI
30.2%
–
60 GPT-5.4 mini (high reasoning)OpenAI
29.4%
$0.75 / $4.50
61 GPT-4o (2024-08-06)OpenAI
26.0%
$2.50 / $10
62 Qwen3.5-Plus (2026-02-15)Alibaba
25.4%
$0.26 / $1.56
63 GPT-5 mini (high reasoning)OpenAI
21.6%
$0.25 / $2
64 Qwen3.5-FlashAlibaba
20.3%
$0.065 / $0.26
65 o4-mini (low reasoning)OpenAI · best of 2 settings
19.6%
$1.10 / $4.40
66 Inkling Small (extra-high reasoning)Thinking Machines
19.1%
$0.45 / $1.20
67 Qwen3.6-FlashAlibaba
15.9%
$0.19 / $1.12
68 o3-mini (high reasoning)OpenAI
15.3%
$1.10 / $4.40
69 Claude Haiku 4.5Anthropic · best of 2 settings
13.2%
$1 / $5
70 GPT-4.1 miniOpenAI
12.7%
$0.40 / $1.60
71 Claude 3 OpusAnthropic
12.6%
–
72 GPT-5.4 nano (high reasoning)OpenAI
11.7%
$0.20 / $1.25
72 GPT-5 nano (high reasoning)OpenAI
11.7%
$0.05 / $0.40
74 Gemma 4 31BGoogle
10.4%
$0.14 / $0.40
75 GPT-4o mini (2024-07-18)OpenAI
8.3%
$0.15 / $0.60
76 GPT-4.1 nanoOpenAI
6.0%
$0.10 / $0.40

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by SimpleQA Verified (Epoch AI); we do not re-run them.

What it measures

Share of short factual questions answered correctly without search, as run by Epoch AI.

What it does not measure

Not answers grounded in your documents; tests what the model remembers.

Epoch AI, Capabilities & benchmarking (CC BY 4.0). Licence: Creative Commons Attribution 4.0 International.