Benchmarks / APEX-Agents

Reported by APEX-Agents

APEX-Agents

Share of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).

Results dated
1 Oct 2026
Models
39
Unit
% of tasks
Licence
Published with permission (Mercor results via Epoch AI)

Full results

APEX-Agents: tasks passed, % of tasks, higher is better
#ModelTasks passed
% of tasks, higher is better
1 Claude Sonnet 5.5 (max)Anthropic
75.5%
2 Claude Opus 5.5 (max reasoning)Anthropic
73.5%
3 Claude Fable 5.1Anthropic
68.6%
4 Gemini 3.7 FlashGoogle
67.8%
5 Claude Opus 5 (max)Anthropic
65.8%
6 Grok 4.6xAI
65.3%
7 GPT-6 AstraOpenAI
64.7%
8 Gemini 3.8 FlashGoogle
64.3%
9 Claude Fable 5Anthropic
63.6%
10 GPT-6.1 Sol (max)OpenAI
60.0%
11 Claude Fable 5.1 (high reasoning)Anthropic
59.7%
12 GPT-5.6 Terra (max)OpenAI
58.2%
13 Muse Spark 1.3Meta
57.8%
14 GLM 5.3Z.ai
56.6%
15 Grok 4.5xAI
56.2%
16 GPT-5.5OpenAI
55.1%
17 Claude Sonnet 5Anthropic
54.5%
18 GPT-6 Sol (max)OpenAI
54.3%
19 GLM 5.3 FlashZ.ai
52.8%
20 GPT-5.4OpenAI
52.4%
21 GPT-5.6 Sol Pro (max)OpenAI
51.4%
22 Kimi K3Moonshot AI
50.6%
23 Claude Opus 4.7 (max reasoning)Anthropic
49.2%
24 Claude Opus 4.8 (max reasoning)Anthropic
48.9%
25 DeepSeek V4 Pro 0813DeepSeek
47.3%
26 Gemini 3.6 FlashGoogle
46.9%
27 Claude Opus 4.6 (max reasoning)Anthropic
46.3%
28 Claude Sonnet 5.5 (medium)Anthropic
44.6%
29 GPT-6 Luna (max)OpenAI
44.3%
30 Claude Sonnet 4.6 (high reasoning)Anthropic
43.0%
31 GLM 5.1Z.ai
40.9%
32 MiniMax M3MiniMax
37.7%
33 Kimi K2.7 CodeMoonshot AI
37.6%
34 Muse Spark 1.2Meta
36.4%
35 Gemini 3.1 Pro PreviewGoogle
35.3%
36 InklingThinkingmachines
33.8%
37 Gemini 3.5 FlashGoogle
27.5%
38 Qwen3.5 397B A17BAlibaba
24.9%
39 gpt-oss-120bOpenAI
4.4%

Results as published by APEX-Agents; we do not re-run them.

What it measures

Share of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).

What it does not measure

Not your firm's documents or tools; graded by rubric, not by a client.

Mercor; collected by Epoch AI. Licence: Published with permission (Mercor results via Epoch AI).