Benchmarks / APEX-Agents
Reported by APEX-Agents
APEX-Agents
Share of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).
- Results dated
- 1 Oct 2026
- Models
- 39
- Unit
- % of tasks
Full results
| # | Model | Tasks passed % of tasks, higher is better |
|---|---|---|
| 1 | Claude Sonnet 5.5 (max)Anthropic |
75.5%
|
| 2 | Claude Opus 5.5 (max reasoning)Anthropic |
73.5%
|
| 3 | Claude Fable 5.1Anthropic |
68.6%
|
| 4 | Gemini 3.7 FlashGoogle |
67.8%
|
| 5 | Claude Opus 5 (max)Anthropic |
65.8%
|
| 6 | Grok 4.6xAI |
65.3%
|
| 7 | GPT-6 AstraOpenAI |
64.7%
|
| 8 | Gemini 3.8 FlashGoogle |
64.3%
|
| 9 | Claude Fable 5Anthropic |
63.6%
|
| 10 | GPT-6.1 Sol (max)OpenAI |
60.0%
|
| 11 | Claude Fable 5.1 (high reasoning)Anthropic |
59.7%
|
| 12 | GPT-5.6 Terra (max)OpenAI |
58.2%
|
| 13 | Muse Spark 1.3Meta |
57.8%
|
| 14 | GLM 5.3Z.ai |
56.6%
|
| 15 | Grok 4.5xAI |
56.2%
|
| 16 | GPT-5.5OpenAI |
55.1%
|
| 17 | Claude Sonnet 5Anthropic |
54.5%
|
| 18 | GPT-6 Sol (max)OpenAI |
54.3%
|
| 19 | GLM 5.3 FlashZ.ai |
52.8%
|
| 20 | GPT-5.4OpenAI |
52.4%
|
| 21 | GPT-5.6 Sol Pro (max)OpenAI |
51.4%
|
| 22 | Kimi K3Moonshot AI |
50.6%
|
| 23 | Claude Opus 4.7 (max reasoning)Anthropic |
49.2%
|
| 24 | Claude Opus 4.8 (max reasoning)Anthropic |
48.9%
|
| 25 | DeepSeek V4 Pro 0813DeepSeek |
47.3%
|
| 26 | Gemini 3.6 FlashGoogle |
46.9%
|
| 27 | Claude Opus 4.6 (max reasoning)Anthropic |
46.3%
|
| 28 | Claude Sonnet 5.5 (medium)Anthropic |
44.6%
|
| 29 | GPT-6 Luna (max)OpenAI |
44.3%
|
| 30 | Claude Sonnet 4.6 (high reasoning)Anthropic |
43.0%
|
| 31 | GLM 5.1Z.ai |
40.9%
|
| 32 | MiniMax M3MiniMax |
37.7%
|
| 33 | Kimi K2.7 CodeMoonshot AI |
37.6%
|
| 34 | Muse Spark 1.2Meta |
36.4%
|
| 35 | Gemini 3.1 Pro PreviewGoogle |
35.3%
|
| 36 | InklingThinkingmachines |
33.8%
|
| 37 | Gemini 3.5 FlashGoogle |
27.5%
|
| 38 | Qwen3.5 397B A17BAlibaba |
24.9%
|
| 39 | gpt-oss-120bOpenAI |
4.4%
|
Results as published by APEX-Agents; we do not re-run them.
What it measures
Share of investment banking, consulting and corporate law tasks an agent completes in a simulated workplace with files and apps, graded against expert criteria (one attempt).
What it does not measure
Not your firm's documents or tools; graded by rubric, not by a client.
Source
Mercor; collected by Epoch AI. Licence: Published with permission (Mercor results via Epoch AI).