Benchmarks / OpenHands Index

Reported by OpenHands Index

OpenHands Index

Answering questions that need web browsing and tools.

Results dated
30 Jun 2026
Models
34
Unit
% resolved
Licence
Apache License 2.0
OpenHands Index: information gathering (GAIA), % resolved, higher is better
#ModelOpenHands Index: information gathering (GAIA)
% resolved, higher is better
1 GPT-5.5OpenAI · gpt-5.5
86.1%
2 Claude Fable 5Anthropic · claude-fable-5
84.2%
3 GPT-5.4OpenAI · gpt-5.4
82.4%
4 Gemini 3.1 Pro PreviewGoogle · gemini-3.1-pro-preview
81.8%
5 Claude Opus 4.7Anthropic · claude-opus-4.7
81.2%
5 Gemini 3.5 FlashGoogle · gemini-3.5-flash
81.2%
7 Claude Opus 4.6Anthropic · claude-opus-4.6
80.0%
8 Claude Opus 4.8Anthropic · claude-opus-4.8
78.8%
9 Kimi K2.6Moonshot AI · kimi-k2.6
74.5%
10 Claude Sonnet 4.5Anthropic · claude-sonnet-4.5
72.7%
11 Qwen3.6 PlusAlibaba · qwen3.6-plus
72.1%
12 GPT-5.2-CodexOpenAI · gpt-5.2-codex
70.9%
13 Claude Opus 4.5Anthropic · claude-opus-4.5
69.1%
14 GLM 5.1Z.ai · glm-5.1
67.3%
15 MiniMax M3MiniMax · minimax-m3
66.7%
16 GPT-5.2OpenAI · gpt-5.2
65.5%
17 Kimi K2.5Moonshot AI · kimi-k2.5
63.6%
18 GLM 5Z.ai · glm-5
60.0%
19 gemini-3-flashGoogle
58.8%
20 GLM 4.7Z.ai · glm-4.7
53.9%
21 Qwen3 Coder NextAlibaba · qwen3-coder-next
50.9%
22 DeepSeek V3.2DeepSeek · deepseek-v3.2
50.3%
23 Qwen3.5-FlashAlibaba · qwen3.5-flash-02-23
49.7%
24 MiniMax M2.5MiniMax · minimax-m2.5
47.9%
25 gemini-3-proGoogle
44.2%
26 Kimi K2 ThinkingMoonshot AI · kimi-k2-thinking
43.6%
27 MiniMax M2.1MiniMax · minimax-m2.1
40.6%
28 Nemotron 3 SuperNVIDIA · nemotron-3-super-120b-a12b
40.0%
29 Qwen3 Coder 480B A35BAlibaba · qwen3-coder
33.9%
30 Trinity Large ThinkingArcee Ai · trinity-large-thinking
32.7%
31 MiniMax M2.7MiniMax · minimax-m2.7
25.5%
32 Claude Sonnet 4.6Anthropic · claude-sonnet-4.6
13.3%
33 DeepSeek V4 Pro 0423DeepSeek · deepseek-v4-pro
12.7%
34 Nemotron 3 Nano 30B A3BNVIDIA · nemotron-3-nano-30b-a3b
8.5%

Results as published by OpenHands Index; we do not re-run them.

What it measures

Answering questions that need web browsing and tools.

What it does not measure

Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.

OpenHands Index by the OpenHands contributors, Apache License 2.0.