Benchmarks / Vending-Bench 2

Reported by Vending-Bench 2

Vending-Bench 2

Bank balance after an agent runs a simulated vending business for a year: ordering stock, setting prices and dealing with suppliers.

Results dated
1 Oct 2026
Models
63
Unit
US dollars
Licence
Published with permission (Andon Labs results via Epoch AI)

Full results

Vending-Bench 2: money after a year, US dollars, higher is better
#ModelMoney after a year
US dollars, higher is better
1 GPT-6 AstraOpenAI
$15,514.70
2 GPT-6 SolOpenAI
$14,427.85
3 Claude Opus 5Anthropic
$11,181.87
4 Claude Opus 4.7Anthropic
$10,936.76
5 GPT-5.6 SolOpenAI
$9,619.37
6 Grok 4.6xAI
$9,047.03
7 GLM 5.2Z.ai
$8,313.78
8 GLM 5.3Z.ai
$8,163.61
9 Claude Opus 4.6Anthropic
$8,017.59
10 GPT-5.5OpenAI
$7,523.84
11 GPT-5.6 TerraOpenAI
$7,343.21
12 Claude Sonnet 4.6Anthropic
$7,204.14
13 Muse Spark 1.1Meta
$6,520.48
14 Claude Sonnet 5Anthropic
$6,377.70
15 Kimi K2.6Moonshot AI
$6,204.57
16 GPT-5.4OpenAI
$6,144.18
17 GPT-5.3-CodexOpenAI
$5,940.12
18 Claude Opus 4.8Anthropic
$5,787.43
19 Claude Fable 5 (high)Anthropic
$5,680.26
20 GLM 5.1Z.ai
$5,634.41
21 gemini-3-pro-previewGoogle
$5,478.16
22 Claude Fable 5.1Anthropic
$5,421.56
23 Gemini 3.5 FlashGoogle
$5,396.42
24 Kimi K3Moonshot AI
$5,165.03
25 Qwen3.6 PlusAlibaba
$5,114.87
26 Gemini 3.8 FlashGoogle
$5,093.79
27 Kimi K2.7 CodeMoonshot AI
$5,082.94
28 Claude Fable 5 (low reasoning)Anthropic
$5,018.52
29 Claude Opus 4.5Anthropic
$4,967.06
30 Claude Fable 5 (max reasoning)Anthropic
$4,966.64
31 Grok 4.20xAI
$4,662.85
32 Claude Fable 5Anthropic
$4,529.94
33 GLM 5Z.ai
$4,432.12
34 Claude Fable 5 (medium reasoning)Anthropic
$4,339.81
35 Qwen3.6 Max PreviewAlibaba
$4,254.19
36 GPT-5.6 LunaOpenAI
$4,094.71
37 Grok 4.5xAI
$3,887.43
38 Claude Sonnet 4.5Anthropic
$3,838.74
39 Gemini 3.1 Pro Preview Custom ToolsGoogle
$3,774.25
40 Gemini 3 Flash PreviewGoogle
$3,634.72
41 GPT-5.2OpenAI
$3,591.33
42 DeepSeek V4 Pro 0423DeepSeek
$3,284.52
43 Claude Opus 4.8 (max reasoning)Anthropic
$2,992.34
44 GLM 4.7Z.ai
$2,376.82
45 MiniMax M3MiniMax
$2,157.77
46 GPT-5.1OpenAI
$1,473.43
47 Kimi K2.5Moonshot AI
$1,198.46
48 grok-4-1-fast-reasoningxAI
$1,106.63
49 DeepSeek V3.2 ExpDeepSeek
$1,034.00
50 Gemini 3.1 Pro PreviewGoogle
$911.21
51 Gemini 2.5 ProGoogle
$573.64
52 Gemini 2.5 FlashGoogle
$548.84
53 Qwen3.5-FlashAlibaba
$462.69
54 Claude Haiku 4.5Anthropic
$458.89
55 Qwen3.5-27BAlibaba
$201.98
56 MiniMax M2MiniMax
$160.60
57 Qwen3 MaxAlibaba
$71.56
58 Grok 4.3xAI
$35.26
59 Qwen3.5 Plus 2026-02-15Alibaba
$0.54
60 Qwen3 235B A22B Thinking 2507Alibaba
$-11.34
61 gpt-oss-120bOpenAI
$-21.53
62 MiniMax M2.5MiniMax
$-23.16
63 GPT-5 MiniOpenAI
$-31.18

Results as published by Vending-Bench 2; we do not re-run them.

What it measures

Bank balance after an agent runs a simulated vending business for a year: ordering stock, setting prices and dealing with suppliers.

What it does not measure

Not a real business; one simulated market with set rules.

Andon Labs; collected by Epoch AI. Licence: Published with permission (Andon Labs results via Epoch AI).