Confirm Action

Are you sure you want to proceed?

Attributed direct benchmark · exact agent configurations

Best OpenHands configurations for repository issue resolution

Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.

OpenHands v1.28.0 + claude-fable-5 leads the official point order at 95.8% resolved.

Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.

Source configurations

34

Task-bootstrap interval available

29

Interval unavailable

5

Official source-score point order

Each subject is the displayed OpenHands version plus language model. Equal scores receive the same competition rank. The deterministic point-order position only arranges tied rows for display; it is not a strict performance claim.

Published 2026-07-20

Point order Exact OpenHands configuration Official score 95% task-bootstrap interval Evidence status
#1 OpenHands v1.28.0 + claude-fable-5 Source identity: OpenHands/claude-fable-5 OpenHands agent v1.28.0 · vision enabled 95.8%SWE-Bench resolved 94.0–97.4%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#2 OpenHands v1.24.0 + claude-opus-4-8 Source identity: OpenHands/claude-opus-4-8 OpenHands agent v1.24.0 · vision enabled 83.8%SWE-Bench resolved 80.6–87.0%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#3 OpenHands v1.24.0 + claude-opus-4-7 Source identity: OpenHands/claude-opus-4-7 OpenHands agent v1.24.0 · vision enabled 81.6%SWE-Bench resolved 78.2–85.0%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#4 OpenHands v1.28.0 + Gemini-3.5-Flash Source identity: OpenHands/Gemini-3.5-Flash OpenHands agent v1.28.0 · vision enabled 78.6%SWE-Bench resolved 75.0–82.2%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#5 OpenHands v1.18.1 + GPT-5.5 Source identity: OpenHands/GPT-5.5 OpenHands agent v1.18.1 · vision enabled 78.2%SWE-Bench resolved 74.6–81.8%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#6tie · display position 6 OpenHands v1.27.0 + Gemini-3.1-Pro Source identity: OpenHands/Gemini-3.1-Pro OpenHands agent v1.27.0 · vision enabled 76.8%SWE-Bench resolved Unavailableofficial point score only Interval unavailable

The pinned task sidecar does not exactly reconstruct the official source headline score, so a task-bootstrap interval is unavailable. The official source headline score remains authoritative.

#6tie · display position 7 OpenHands v1.15.0 + claude-opus-4-6 Source identity: OpenHands/claude-opus-4-6 OpenHands agent v1.15.0 · vision enabled 76.8%SWE-Bench resolved 73.2–80.4%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#8 OpenHands v1.8.3 + claude-opus-4-5 Source identity: OpenHands/claude-opus-4-5 OpenHands agent v1.8.3 · vision enabled 76.6%SWE-Bench resolved 73.0–80.2%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#9 OpenHands v1.24.0 + MiniMax-M3 Source identity: OpenHands/MiniMax-M3 OpenHands agent v1.24.0 · vision enabled 76.4%SWE-Bench resolved 72.6–80.2%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#10tie · display position 10 OpenHands v1.18.1 + GPT-5.4 Source identity: OpenHands/GPT-5.4 OpenHands agent v1.18.1 · vision enabled 75.6%SWE-Bench resolved 71.8–79.4%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#10tie · display position 11 OpenHands v1.14.0 + Minimax-2.7 Source identity: OpenHands/MiniMax-M2.7 OpenHands agent v1.14.0 75.6%SWE-Bench resolved 71.8–79.2%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#12 OpenHands v1.17.0 + GLM-5.1 Source identity: OpenHands/GLM-5.1 OpenHands agent v1.17.0 75.0%SWE-Bench resolved 71.2–78.8%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#13tie · display position 13 OpenHands v1.8.3 + GPT-5.2 Source identity: OpenHands/GPT-5.2 OpenHands agent v1.8.3 · vision enabled 74.6%SWE-Bench resolved 70.8–78.4%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#13tie · display position 14 OpenHands v1.8.3 + Gemini-3-Flash Source identity: OpenHands/Gemini-3-Flash OpenHands agent v1.8.3 · vision enabled 74.6%SWE-Bench resolved 70.8–78.4%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#13tie · display position 15 OpenHands v1.18.1 + Kimi-K2.6 Source identity: OpenHands/Kimi-K2.6 OpenHands agent v1.18.1 · vision enabled 74.6%SWE-Bench resolved 71.0–78.4%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#16 OpenHands v1.11.5 + claude-sonnet-4-6 Source identity: OpenHands/claude-sonnet-4-6 OpenHands agent v1.11.5 · vision enabled 74.4%SWE-Bench resolved 70.6–78.2%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#17tie · display position 17 OpenHands v1.16.1 + Qwen3.6-Plus Source identity: OpenHands/Qwen3.6-Plus OpenHands agent v1.16.1 · vision enabled 74.2%SWE-Bench resolved 70.4–78.0%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#17tie · display position 18 OpenHands v1.8.3 + claude-sonnet-4-5 Source identity: OpenHands/claude-sonnet-4-5 OpenHands agent v1.8.3 · vision enabled 74.2%SWE-Bench resolved 70.4–78.0%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#19 OpenHands v1.8.3 + GPT-5.2-Codex Source identity: OpenHands/GPT-5.2-Codex OpenHands agent v1.8.3 · vision enabled 73.8%SWE-Bench resolved 70.0–77.6%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#20tie · display position 20 OpenHands v1.10.0 + GLM-4.7 Source identity: OpenHands/GLM-4.7 OpenHands agent v1.10.0 73.4%SWE-Bench resolved 69.6–77.2%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#20tie · display position 21 OpenHands v1.11.5 + GLM-5 Source identity: OpenHands/GLM-5 OpenHands agent v1.11.5 73.4%SWE-Bench resolved 69.6–77.4%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#22 OpenHands v1.22.1 + DeepSeek-V4-Pro Source identity: OpenHands/DeepSeek-V4-Pro OpenHands agent v1.22.1 73.2%SWE-Bench resolved 69.2–77.2%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#23 OpenHands v1.11.3 + MiniMax-M2.5 Source identity: OpenHands/MiniMax-M2.5 OpenHands agent v1.11.3 72.6%SWE-Bench resolved 68.6–76.4%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#24 OpenHands v1.8.3 + DeepSeek-V3.2-Reasoner Source identity: OpenHands/DeepSeek-V3.2-Reasoner OpenHands agent v1.8.3 71.6%SWE-Bench resolved Unavailableofficial point score only Interval unavailable

The pinned task sidecar contains one or more unknown outcomes, so a task-bootstrap interval is unavailable. The official source headline score is retained without treating unknown outcomes as failures.

#25 OpenHands v1.8.3 + Gemini-3-Pro Source identity: OpenHands/Gemini-3-Pro OpenHands agent v1.8.3 · vision enabled 70.6%SWE-Bench resolved 66.6–74.6%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#26 OpenHands v1.8.3 + Kimi-K2-Thinking Source identity: OpenHands/Kimi-K2-Thinking OpenHands agent v1.8.3 69.2%SWE-Bench resolved Unavailableofficial point score only Interval unavailable

The pinned task sidecar contains one or more unknown outcomes, so a task-bootstrap interval is unavailable. The official source headline score is retained without treating unknown outcomes as failures.

#27tie · display position 27 OpenHands v1.8.3 + Kimi-K2.5 Source identity: OpenHands/Kimi-K2.5 OpenHands agent v1.8.3 · vision enabled 68.8%SWE-Bench resolved Unavailableofficial point score only Interval unavailable

The pinned task sidecar contains one or more unknown outcomes, so a task-bootstrap interval is unavailable. The official source headline score is retained without treating unknown outcomes as failures.

#27tie · display position 28 OpenHands v1.8.3 + MiniMax-M2.1 Source identity: OpenHands/MiniMax-M2.1 OpenHands agent v1.8.3 68.8%SWE-Bench resolved 64.6–72.8%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#29 OpenHands v1.11.1 + Qwen3-Coder-Next Source identity: OpenHands/Qwen3-Coder-Next OpenHands agent v1.11.1 66.6%SWE-Bench resolved 62.4–70.8%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#30 OpenHands v1.8.3 + Qwen3-Coder-480B Source identity: OpenHands/Qwen3-Coder-480B OpenHands agent v1.8.3 62.4%SWE-Bench resolved 58.2–66.6%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#31tie · display position 31 OpenHands v1.16.1 + Nemotron-3-Super Source identity: OpenHands/Nemotron-3-Super OpenHands agent v1.16.1 62.0%SWE-Bench resolved 57.8–66.2%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#31tie · display position 32 OpenHands v1.16.1 + Qwen3.5-Flash Source identity: OpenHands/Qwen3.5-Flash OpenHands agent v1.16.1 · vision enabled 62.0%SWE-Bench resolved 57.8–66.4%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#33 OpenHands v1.16.1 + Trinity-Large-Thinking Source identity: OpenHands/Trinity-Large-Thinking OpenHands agent v1.16.1 56.8%SWE-Bench resolved 52.4–61.2%marginal score interval Exact 500-task reconstruction

Sidecar exactly reproduces the official headline. No rank-confidence claim.

#34 OpenHands v1.8.3 + Nemotron-3-Nano Source identity: OpenHands/Nemotron-3-Nano OpenHands agent v1.8.3 34.2%SWE-Bench resolved Unavailableofficial point score only Interval unavailable

The pinned task sidecar contains one or more unknown outcomes, so a task-bootstrap interval is unavailable. The official source headline score is retained without treating unknown outcomes as failures.

No Artificial Analysis price or throughput data is joined here: the benchmark subjects are exact OpenHands configurations, and no reviewed configuration-level operational match exists. Read the full source methodology →

What the score means

Official OpenHands SWE-Bench resolved percentage for the exact OpenHands and language-model configuration. It measures the full agent configuration, not the language model in isolation.

The five interval-unavailable rows keep the publisher's official score. Missing outcomes are not converted to failures, and divergent sidecars are not used to replace the official headline.

Not part of the overall leaderboard

The publisher-native resolved percentage measures an OpenHands configuration on SWE-Bench. It is neither a calibrated SpringPrompt task-fit score nor a bare-model score, so cross-task scale compatibility has not been established.

Scale compatibility: not-proven

Frequently asked

Is this a ranking of the bare language models?

No. Each row is the exact OpenHands agent version plus its source-listed language model. A different agent, tool setup or version can produce a different result.

Why do five rows have no score interval?

Their official OpenHands point scores remain available, but the pinned task sidecars either contain unknown outcomes or do not exactly reproduce the official headline. We do not repair or coerce those records.

Does a point position represent confident rank?

No. The table sorts official point scores and uses competition ranks for ties. Because five rows lack comparable task-level uncertainty, no full-cohort rank-confidence claim is made.

Why are there no cheapest or fastest views?

The evaluated subjects are exact OpenHands configurations. We have not verified an exact configuration-level operational join, so bare-model price or throughput is not attached to these rows.

Does this enter SpringPrompt's overall leaderboard?

No. The publisher-native resolved percentage measures an OpenHands configuration on SWE-Bench. It is neither a calibrated SpringPrompt task-fit score nor a bare-model score, so cross-task scale compatibility has not been established.