Benchmarks / OpenHands Index
Reported by OpenHands Index
OpenHands Index
Resolving real GitHub issues so that the repository's own tests pass.
- Results dated
- 30 Jun 2026
- Models
- 34
- Unit
- % resolved
- Licence
- Apache License 2.0
| # | Model | OpenHands Index: issue resolution (SWE-Bench) % resolved, higher is better |
|---|---|---|
| 1 | Claude Fable 5Anthropic · claude-fable-5 |
95.8%
|
| 2 | Claude Opus 4.8Anthropic · claude-opus-4.8 |
83.8%
|
| 3 | Claude Opus 4.7Anthropic · claude-opus-4.7 |
81.6%
|
| 4 | Gemini 3.5 FlashGoogle · gemini-3.5-flash |
78.6%
|
| 5 | GPT-5.5OpenAI · gpt-5.5 |
78.2%
|
| 6 | Claude Opus 4.6Anthropic · claude-opus-4.6 |
76.8%
|
| 6 | Gemini 3.1 Pro PreviewGoogle · gemini-3.1-pro-preview |
76.8%
|
| 8 | Claude Opus 4.5Anthropic · claude-opus-4.5 |
76.6%
|
| 9 | MiniMax M3MiniMax · minimax-m3 |
76.4%
|
| 10 | MiniMax M2.7MiniMax · minimax-m2.7 |
75.6%
|
| 10 | GPT-5.4OpenAI · gpt-5.4 |
75.6%
|
| 12 | GLM 5.1Z.ai · glm-5.1 |
75.0%
|
| 13 | gemini-3-flashGoogle |
74.6%
|
| 13 | Kimi K2.6Moonshot AI · kimi-k2.6 |
74.6%
|
| 13 | GPT-5.2OpenAI · gpt-5.2 |
74.6%
|
| 16 | Claude Sonnet 4.6Anthropic · claude-sonnet-4.6 |
74.4%
|
| 17 | Qwen3.6 PlusAlibaba · qwen3.6-plus |
74.2%
|
| 17 | Claude Sonnet 4.5Anthropic · claude-sonnet-4.5 |
74.2%
|
| 19 | GPT-5.2-CodexOpenAI · gpt-5.2-codex |
73.8%
|
| 20 | GLM 4.7Z.ai · glm-4.7 |
73.4%
|
| 20 | GLM 5Z.ai · glm-5 |
73.4%
|
| 22 | DeepSeek V4 Pro 0423DeepSeek · deepseek-v4-pro |
73.2%
|
| 23 | MiniMax M2.5MiniMax · minimax-m2.5 |
72.6%
|
| 24 | DeepSeek V3.2DeepSeek · deepseek-v3.2 |
71.6%
|
| 25 | gemini-3-proGoogle |
70.6%
|
| 26 | Kimi K2 ThinkingMoonshot AI · kimi-k2-thinking |
69.2%
|
| 27 | MiniMax M2.1MiniMax · minimax-m2.1 |
68.8%
|
| 27 | Kimi K2.5Moonshot AI · kimi-k2.5 |
68.8%
|
| 29 | Qwen3 Coder NextAlibaba · qwen3-coder-next |
66.6%
|
| 30 | Qwen3 Coder 480B A35BAlibaba · qwen3-coder |
62.4%
|
| 31 | Qwen3.5-FlashAlibaba · qwen3.5-flash-02-23 |
62.0%
|
| 31 | Nemotron 3 SuperNVIDIA · nemotron-3-super-120b-a12b |
62.0%
|
| 33 | Trinity Large ThinkingArcee Ai · trinity-large-thinking |
56.8%
|
| 34 | Nemotron 3 Nano 30B A3BNVIDIA · nemotron-3-nano-30b-a3b |
34.2%
|
Results as published by OpenHands Index; we do not re-run them.
What it measures
Resolving real GitHub issues so that the repository's own tests pass.
What it does not measure
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Source
OpenHands Index by the OpenHands contributors, Apache License 2.0.