Attributed direct benchmark · exact agent configurations
Best OpenHands configurations for repository issue resolution
Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.
OpenHands v1.28.0 + claude-fable-5 leads the official point order at 95.8% resolved.
Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.
Source configurations
34
Task-bootstrap interval available
29
Interval unavailable
5
Official source-score point order
Each subject is the displayed OpenHands version plus language model. Equal scores receive the same competition rank. The deterministic point-order position only arranges tied rows for display; it is not a strict performance claim.
Published 2026-07-20
| Point order | Exact OpenHands configuration | Official score | 95% task-bootstrap interval | Evidence status |
|---|---|---|---|---|
| #1 | OpenHands v1.28.0 + claude-fable-5 Source identity: OpenHands/claude-fable-5 OpenHands agent v1.28.0 · vision enabled | 95.8%SWE-Bench resolved | 94.0–97.4%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #2 | OpenHands v1.24.0 + claude-opus-4-8 Source identity: OpenHands/claude-opus-4-8 OpenHands agent v1.24.0 · vision enabled | 83.8%SWE-Bench resolved | 80.6–87.0%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #3 | OpenHands v1.24.0 + claude-opus-4-7 Source identity: OpenHands/claude-opus-4-7 OpenHands agent v1.24.0 · vision enabled | 81.6%SWE-Bench resolved | 78.2–85.0%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #4 | OpenHands v1.28.0 + Gemini-3.5-Flash Source identity: OpenHands/Gemini-3.5-Flash OpenHands agent v1.28.0 · vision enabled | 78.6%SWE-Bench resolved | 75.0–82.2%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #5 | OpenHands v1.18.1 + GPT-5.5 Source identity: OpenHands/GPT-5.5 OpenHands agent v1.18.1 · vision enabled | 78.2%SWE-Bench resolved | 74.6–81.8%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #6tie · display position 6 | OpenHands v1.27.0 + Gemini-3.1-Pro Source identity: OpenHands/Gemini-3.1-Pro OpenHands agent v1.27.0 · vision enabled | 76.8%SWE-Bench resolved | Unavailableofficial point score only |
Interval unavailable The pinned task sidecar does not exactly reconstruct the official source headline score, so a task-bootstrap interval is unavailable. The official source headline score remains authoritative. |
| #6tie · display position 7 | OpenHands v1.15.0 + claude-opus-4-6 Source identity: OpenHands/claude-opus-4-6 OpenHands agent v1.15.0 · vision enabled | 76.8%SWE-Bench resolved | 73.2–80.4%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #8 | OpenHands v1.8.3 + claude-opus-4-5 Source identity: OpenHands/claude-opus-4-5 OpenHands agent v1.8.3 · vision enabled | 76.6%SWE-Bench resolved | 73.0–80.2%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #9 | OpenHands v1.24.0 + MiniMax-M3 Source identity: OpenHands/MiniMax-M3 OpenHands agent v1.24.0 · vision enabled | 76.4%SWE-Bench resolved | 72.6–80.2%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #10tie · display position 10 | OpenHands v1.18.1 + GPT-5.4 Source identity: OpenHands/GPT-5.4 OpenHands agent v1.18.1 · vision enabled | 75.6%SWE-Bench resolved | 71.8–79.4%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #10tie · display position 11 | OpenHands v1.14.0 + Minimax-2.7 Source identity: OpenHands/MiniMax-M2.7 OpenHands agent v1.14.0 | 75.6%SWE-Bench resolved | 71.8–79.2%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #12 | OpenHands v1.17.0 + GLM-5.1 Source identity: OpenHands/GLM-5.1 OpenHands agent v1.17.0 | 75.0%SWE-Bench resolved | 71.2–78.8%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #13tie · display position 13 | OpenHands v1.8.3 + GPT-5.2 Source identity: OpenHands/GPT-5.2 OpenHands agent v1.8.3 · vision enabled | 74.6%SWE-Bench resolved | 70.8–78.4%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #13tie · display position 14 | OpenHands v1.8.3 + Gemini-3-Flash Source identity: OpenHands/Gemini-3-Flash OpenHands agent v1.8.3 · vision enabled | 74.6%SWE-Bench resolved | 70.8–78.4%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #13tie · display position 15 | OpenHands v1.18.1 + Kimi-K2.6 Source identity: OpenHands/Kimi-K2.6 OpenHands agent v1.18.1 · vision enabled | 74.6%SWE-Bench resolved | 71.0–78.4%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #16 | OpenHands v1.11.5 + claude-sonnet-4-6 Source identity: OpenHands/claude-sonnet-4-6 OpenHands agent v1.11.5 · vision enabled | 74.4%SWE-Bench resolved | 70.6–78.2%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #17tie · display position 17 | OpenHands v1.16.1 + Qwen3.6-Plus Source identity: OpenHands/Qwen3.6-Plus OpenHands agent v1.16.1 · vision enabled | 74.2%SWE-Bench resolved | 70.4–78.0%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #17tie · display position 18 | OpenHands v1.8.3 + claude-sonnet-4-5 Source identity: OpenHands/claude-sonnet-4-5 OpenHands agent v1.8.3 · vision enabled | 74.2%SWE-Bench resolved | 70.4–78.0%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #19 | OpenHands v1.8.3 + GPT-5.2-Codex Source identity: OpenHands/GPT-5.2-Codex OpenHands agent v1.8.3 · vision enabled | 73.8%SWE-Bench resolved | 70.0–77.6%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #20tie · display position 20 | OpenHands v1.10.0 + GLM-4.7 Source identity: OpenHands/GLM-4.7 OpenHands agent v1.10.0 | 73.4%SWE-Bench resolved | 69.6–77.2%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #20tie · display position 21 | OpenHands v1.11.5 + GLM-5 Source identity: OpenHands/GLM-5 OpenHands agent v1.11.5 | 73.4%SWE-Bench resolved | 69.6–77.4%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #22 | OpenHands v1.22.1 + DeepSeek-V4-Pro Source identity: OpenHands/DeepSeek-V4-Pro OpenHands agent v1.22.1 | 73.2%SWE-Bench resolved | 69.2–77.2%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #23 | OpenHands v1.11.3 + MiniMax-M2.5 Source identity: OpenHands/MiniMax-M2.5 OpenHands agent v1.11.3 | 72.6%SWE-Bench resolved | 68.6–76.4%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #24 | OpenHands v1.8.3 + DeepSeek-V3.2-Reasoner Source identity: OpenHands/DeepSeek-V3.2-Reasoner OpenHands agent v1.8.3 | 71.6%SWE-Bench resolved | Unavailableofficial point score only |
Interval unavailable The pinned task sidecar contains one or more unknown outcomes, so a task-bootstrap interval is unavailable. The official source headline score is retained without treating unknown outcomes as failures. |
| #25 | OpenHands v1.8.3 + Gemini-3-Pro Source identity: OpenHands/Gemini-3-Pro OpenHands agent v1.8.3 · vision enabled | 70.6%SWE-Bench resolved | 66.6–74.6%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #26 | OpenHands v1.8.3 + Kimi-K2-Thinking Source identity: OpenHands/Kimi-K2-Thinking OpenHands agent v1.8.3 | 69.2%SWE-Bench resolved | Unavailableofficial point score only |
Interval unavailable The pinned task sidecar contains one or more unknown outcomes, so a task-bootstrap interval is unavailable. The official source headline score is retained without treating unknown outcomes as failures. |
| #27tie · display position 27 | OpenHands v1.8.3 + Kimi-K2.5 Source identity: OpenHands/Kimi-K2.5 OpenHands agent v1.8.3 · vision enabled | 68.8%SWE-Bench resolved | Unavailableofficial point score only |
Interval unavailable The pinned task sidecar contains one or more unknown outcomes, so a task-bootstrap interval is unavailable. The official source headline score is retained without treating unknown outcomes as failures. |
| #27tie · display position 28 | OpenHands v1.8.3 + MiniMax-M2.1 Source identity: OpenHands/MiniMax-M2.1 OpenHands agent v1.8.3 | 68.8%SWE-Bench resolved | 64.6–72.8%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #29 | OpenHands v1.11.1 + Qwen3-Coder-Next Source identity: OpenHands/Qwen3-Coder-Next OpenHands agent v1.11.1 | 66.6%SWE-Bench resolved | 62.4–70.8%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #30 | OpenHands v1.8.3 + Qwen3-Coder-480B Source identity: OpenHands/Qwen3-Coder-480B OpenHands agent v1.8.3 | 62.4%SWE-Bench resolved | 58.2–66.6%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #31tie · display position 31 | OpenHands v1.16.1 + Nemotron-3-Super Source identity: OpenHands/Nemotron-3-Super OpenHands agent v1.16.1 | 62.0%SWE-Bench resolved | 57.8–66.2%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #31tie · display position 32 | OpenHands v1.16.1 + Qwen3.5-Flash Source identity: OpenHands/Qwen3.5-Flash OpenHands agent v1.16.1 · vision enabled | 62.0%SWE-Bench resolved | 57.8–66.4%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #33 | OpenHands v1.16.1 + Trinity-Large-Thinking Source identity: OpenHands/Trinity-Large-Thinking OpenHands agent v1.16.1 | 56.8%SWE-Bench resolved | 52.4–61.2%marginal score interval |
Exact 500-task reconstruction Sidecar exactly reproduces the official headline. No rank-confidence claim. |
| #34 | OpenHands v1.8.3 + Nemotron-3-Nano Source identity: OpenHands/Nemotron-3-Nano OpenHands agent v1.8.3 | 34.2%SWE-Bench resolved | Unavailableofficial point score only |
Interval unavailable The pinned task sidecar contains one or more unknown outcomes, so a task-bootstrap interval is unavailable. The official source headline score is retained without treating unknown outcomes as failures. |
No Artificial Analysis price or throughput data is joined here: the benchmark subjects are exact OpenHands configurations, and no reviewed configuration-level operational match exists. Read the full source methodology →
What the score means
Official OpenHands SWE-Bench resolved percentage for the exact OpenHands and language-model configuration. It measures the full agent configuration, not the language model in isolation.
The five interval-unavailable rows keep the publisher's official score. Missing outcomes are not converted to failures, and divergent sidecars are not used to replace the official headline.
Not part of the overall leaderboard
The publisher-native resolved percentage measures an OpenHands configuration on SWE-Bench. It is neither a calibrated SpringPrompt task-fit score nor a bare-model score, so cross-task scale compatibility has not been established.
Scale compatibility: not-proven
Frequently asked
Is this a ranking of the bare language models?
No. Each row is the exact OpenHands agent version plus its source-listed language model. A different agent, tool setup or version can produce a different result.
Why do five rows have no score interval?
Their official OpenHands point scores remain available, but the pinned task sidecars either contain unknown outcomes or do not exactly reproduce the official headline. We do not repair or coerce those records.
Does a point position represent confident rank?
No. The table sorts official point scores and uses competition ranks for ties. Because five rows lack comparable task-level uncertainty, no full-cohort rank-confidence claim is made.
Why are there no cheapest or fastest views?
The evaluated subjects are exact OpenHands configurations. We have not verified an exact configuration-level operational join, so bare-model price or throughput is not attached to these rows.
Does this enter SpringPrompt's overall leaderboard?
No. The publisher-native resolved percentage measures an OpenHands configuration on SWE-Bench. It is neither a calibrated SpringPrompt task-fit score nor a bare-model score, so cross-task scale compatibility has not been established.