Benchmarks / GDP.pdf
Reported by GDP.pdf
GDP.pdf
Rubric score on 100 real professional tasks that need reasoning over PDFs (finance, legal, insurance, engineering, HR and more).
- Results dated
- 6 Jun 2026 to 1 Oct 2026
- Models
- 49
- Unit
- % of rubric
Full results
| # | Model | Professional document tasks % of rubric, higher is better |
|---|---|---|
| 1 | GPT-5.6 Sol (max)OpenAI |
30.7%
|
| 1 | GPT-5.6 SolOpenAI |
30.7%
|
| 3 | Claude Opus 5.5 (max reasoning)Anthropic |
30.6%
|
| 4 | Claude Fable 5Anthropic |
30.0%
|
| 5 | Claude Fable 5 (max reasoning)Anthropic |
29.8%
|
| 6 | Claude Fable 5.1 (high reasoning)Anthropic |
29.6%
|
| 7 | Claude Fable 5.1 (max)Anthropic |
27.6%
|
| 7 | Muse Spark 1.3 (xhigh)Meta |
27.6%
|
| 7 | Muse Spark 1.3Meta |
27.6%
|
| 10 | GPT-6 Sol (max)OpenAI |
26.4%
|
| 11 | GPT-5.5 (xhigh reasoning)OpenAI |
26.0%
|
| 12 | GPT-5.6 Terra (medium reasoning)OpenAI |
24.7%
|
| 12 | GPT-5.6 TerraOpenAI |
24.7%
|
| 14 | Claude Opus 4.8 (max reasoning)Anthropic |
24.0%
|
| 14 | Claude Opus 5 (max)Anthropic |
24.0%
|
| 16 | Gemini 3.7 Flash (high)Google |
23.8%
|
| 17 | Gemini 3.8 Flash (medium reasoning)Google |
23.4%
|
| 18 | Qwen3.8 Max (0902) (xhigh)Alibaba |
23.2%
|
| 18 | Gemini 3.8 Flash (high)Google |
23.2%
|
| 20 | GPT-6 Luna (max)OpenAI |
23.0%
|
| 21 | GPT-5.6 Luna (medium reasoning)OpenAI |
22.7%
|
| 21 | GPT-5.6 LunaOpenAI |
22.7%
|
| 23 | Gemini 3.7 Flash (medium reasoning)Google |
21.8%
|
| 24 | Claude Opus 4.7 (max reasoning)Anthropic |
21.0%
|
| 25 | Kimi K3 (max)Moonshot AI |
19.0%
|
| 26 | Claude Sonnet 4.6 (max reasoning)Anthropic |
18.0%
|
| 27 | Grok 4.6 (xhigh)xAI |
17.2%
|
| 28 | Gemini 3.1 Pro Preview (high reasoning)Google |
17.0%
|
| 28 | Gemini 3.1 Pro PreviewGoogle |
17.0%
|
| 30 | Qwen3.8 Flash (xhigh)Alibaba |
16.6%
|
| 31 | Muse Spark 1.2Meta |
16.0%
|
| 31 | Grok 4.6 (high)xAI |
16.0%
|
| 33 | Muse Spark 1.1 (medium)Meta |
15.0%
|
| 33 | Muse Spark 1.1Meta |
15.0%
|
| 35 | Gemini 3.5 Flash (medium)Google |
14.0%
|
| 35 | Gemini 3.5 FlashGoogle |
14.0%
|
| 35 | Gemini 3.6 Flash (high)Google |
14.0%
|
| 35 | Gemini 3.6 FlashGoogle |
14.0%
|
| 35 | Grok 4.5 (high)xAI |
14.0%
|
| 35 | GLM 5.3 Flash (max)Z.ai |
14.0%
|
| 41 | Muse Spark 1.2Meta |
12.0%
|
| 41 | Kimi K2.6Moonshot AI |
12.0%
|
| 43 | Gemini 3.5 Flash Lite (high)Google |
10.0%
|
| 43 | Gemini 3 Flash Preview (high reasoning)Google |
10.0%
|
| 43 | Gemini 3 Flash PreviewGoogle |
10.0%
|
| 46 | Grok 4.3 (high)xAI |
8.0%
|
| 46 | Grok 4.3xAI |
8.0%
|
| 48 | Nova 2 Pro (preview, no reasoning)Amazon |
2.0%
|
| 48 | Nova 2 Pro (preview)Amazon |
2.0%
|
Results as published by GDP.pdf; we do not re-run them.
What it measures
Rubric score on 100 real professional tasks that need reasoning over PDFs (finance, legal, insurance, engineering, HR and more).
What it does not measure
Not your documents; 100 tasks, so small differences are noise.
Source
Surge AI; collected by Epoch AI. Licence: Published with permission (Surge AI results via Epoch AI).