Benchmarks / GDP.pdf

Reported by GDP.pdf

GDP.pdf

Rubric score on 100 real professional tasks that need reasoning over PDFs (finance, legal, insurance, engineering, HR and more).

Results dated
6 Jun 2026 to 1 Oct 2026
Models
49
Unit
% of rubric
Licence
Published with permission (Surge AI results via Epoch AI)

Full results

GDP.pdf: professional document tasks, % of rubric, higher is better
#ModelProfessional document tasks
% of rubric, higher is better
1 GPT-5.6 Sol (max)OpenAI
30.7%
1 GPT-5.6 SolOpenAI
30.7%
3 Claude Opus 5.5 (max reasoning)Anthropic
30.6%
4 Claude Fable 5Anthropic
30.0%
5 Claude Fable 5 (max reasoning)Anthropic
29.8%
6 Claude Fable 5.1 (high reasoning)Anthropic
29.6%
7 Claude Fable 5.1 (max)Anthropic
27.6%
7 Muse Spark 1.3 (xhigh)Meta
27.6%
7 Muse Spark 1.3Meta
27.6%
10 GPT-6 Sol (max)OpenAI
26.4%
11 GPT-5.5 (xhigh reasoning)OpenAI
26.0%
12 GPT-5.6 Terra (medium reasoning)OpenAI
24.7%
12 GPT-5.6 TerraOpenAI
24.7%
14 Claude Opus 4.8 (max reasoning)Anthropic
24.0%
14 Claude Opus 5 (max)Anthropic
24.0%
16 Gemini 3.7 Flash (high)Google
23.8%
17 Gemini 3.8 Flash (medium reasoning)Google
23.4%
18 Qwen3.8 Max (0902) (xhigh)Alibaba
23.2%
18 Gemini 3.8 Flash (high)Google
23.2%
20 GPT-6 Luna (max)OpenAI
23.0%
21 GPT-5.6 Luna (medium reasoning)OpenAI
22.7%
21 GPT-5.6 LunaOpenAI
22.7%
23 Gemini 3.7 Flash (medium reasoning)Google
21.8%
24 Claude Opus 4.7 (max reasoning)Anthropic
21.0%
25 Kimi K3 (max)Moonshot AI
19.0%
26 Claude Sonnet 4.6 (max reasoning)Anthropic
18.0%
27 Grok 4.6 (xhigh)xAI
17.2%
28 Gemini 3.1 Pro Preview (high reasoning)Google
17.0%
28 Gemini 3.1 Pro PreviewGoogle
17.0%
30 Qwen3.8 Flash (xhigh)Alibaba
16.6%
31 Muse Spark 1.2Meta
16.0%
31 Grok 4.6 (high)xAI
16.0%
33 Muse Spark 1.1 (medium)Meta
15.0%
33 Muse Spark 1.1Meta
15.0%
35 Gemini 3.5 Flash (medium)Google
14.0%
35 Gemini 3.5 FlashGoogle
14.0%
35 Gemini 3.6 Flash (high)Google
14.0%
35 Gemini 3.6 FlashGoogle
14.0%
35 Grok 4.5 (high)xAI
14.0%
35 GLM 5.3 Flash (max)Z.ai
14.0%
41 Muse Spark 1.2Meta
12.0%
41 Kimi K2.6Moonshot AI
12.0%
43 Gemini 3.5 Flash Lite (high)Google
10.0%
43 Gemini 3 Flash Preview (high reasoning)Google
10.0%
43 Gemini 3 Flash PreviewGoogle
10.0%
46 Grok 4.3 (high)xAI
8.0%
46 Grok 4.3xAI
8.0%
48 Nova 2 Pro (preview, no reasoning)Amazon
2.0%
48 Nova 2 Pro (preview)Amazon
2.0%

Results as published by GDP.pdf; we do not re-run them.

What it measures

Rubric score on 100 real professional tasks that need reasoning over PDFs (finance, legal, insurance, engineering, HR and more).

What it does not measure

Not your documents; 100 tasks, so small differences are noise.

Surge AI; collected by Epoch AI. Licence: Published with permission (Surge AI results via Epoch AI).