Attributed direct benchmark · agent configurations
Best coding agent configurations
Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.
Codex - GPT-5.6 Sol (max) leads the complete v1.2 point order at 61.0.
Rows are source-listed coding-agent and model configurations. Artificial Analysis does not expose agent harness versions, so these are not bare-model ranks or immutable runnable-system identities. Quality, cost and wall time are publisher point estimates; no interval, significance or rank-confidence claim is available.
Source configurations
44
Complete 3/3 cohort
40
Legacy 2/3 · unranked
4
Lowest source-reported benchmark cost per task
Positions use the publisher's mean benchmark cost per task for the same exact agent configuration. Cost changes no quality score. Complete rows without a cost remain visible but unpositioned.
Source snapshot 2026-07-20
| Lowest source-reported benchmark cost per task | Exact agent + model configuration | Index point score | DeepSWE | Terminal-Bench v2 | SWE-Atlas-QnA | Cost/task | Wall time/task |
|---|---|---|---|---|---|---|---|
| #1 quality point order #31 | Cursor CLI - Composer 2.5 Agent: Cursor CLI · Model variant: Composer 2.5 Agent version not exposed by source · source row 7e4d99f22241d821b64798b319e655de | 33.8 3-component source point | 15.9 | 67.1 | 18.3 | $0.08mean benchmark USD | 9.6 minmean agent wall time |
| #2 quality point order #38 | Codex - GPT-5.6 Luna (low) Agent: Codex · Model variant: GPT-5.6 Luna (low) Agent version not exposed by source · source row ba14c764e7104dbf50b2f160ebf14426 | 23.4 3-component source point | 10.3 | 49.6 | 10.2 | $0.21mean benchmark USD | 1.9 minmean agent wall time |
| #3 quality point order #37 | Claude Code - DeepSeek V4 Pro (high) Agent: Claude Code · Model variant: DeepSeek V4 Pro (high) Agent version not exposed by source · source row 19819b29f8069ec81aec31ce3fa68958 | 28.5 3-component source point | 8.6 | 65.9 | 11.0 | $0.27mean benchmark USD | 17.9 minmean agent wall time |
| #4 quality point order #40 | Codex - GPT-5.6 Luna (none) Agent: Codex · Model variant: GPT-5.6 Luna (none) Agent version not exposed by source · source row 824070d2da3d80c2664a07004bf77f36 | 17.9 3-component source point | 6.5 | 37.3 | 9.9 | $0.35mean benchmark USD | 2.5 minmean agent wall time |
| #5 quality point order #39 | Codex - GPT-5.6 Terra (none) Agent: Codex · Model variant: GPT-5.6 Terra (none) Agent version not exposed by source · source row 0a3201d26f6411e9328800c6bfbcc1e2 | 21.9 3-component source point | 13.3 | 39.3 | 13.2 | $0.37mean benchmark USD | 1.8 minmean agent wall time |
| #6 quality point order #27 | Codex - GPT-5.6 Luna (medium) Agent: Codex · Model variant: GPT-5.6 Luna (medium) Agent version not exposed by source · source row 8c17551948c19f8277b123a445a01485 | 38.4 3-component source point | 36.6 | 63.5 | 15.1 | $0.47mean benchmark USD | 3.4 minmean agent wall time |
| #7 quality point order #30 | Codex - GPT-5.6 Terra (low) Agent: Codex · Model variant: GPT-5.6 Terra (low) Agent version not exposed by source · source row deffe38fe77a5fd77c871140402adf8f | 33.9 3-component source point | 29.8 | 57.5 | 14.2 | $0.48mean benchmark USD | 2.8 minmean agent wall time |
| #8 quality point order #32 | Cursor CLI - Composer 2.5 Fast Agent: Cursor CLI · Model variant: Composer 2.5 Fast Agent version not exposed by source · source row 913859e7107291a71a76e2ad0bac704a | 33.8 3-component source point | 15.9 | 67.1 | 18.3 | $0.56mean benchmark USD | 6.8 minmean agent wall time |
| #9 quality point order #22 | Codex - GPT-5.6 Terra (medium) Agent: Codex · Model variant: GPT-5.6 Terra (medium) Agent version not exposed by source · source row 9cd9393327109ca9132fff0834a65880 | 43.9 3-component source point | 45.7 | 69.4 | 16.7 | $0.90mean benchmark USD | 4.3 minmean agent wall time |
| #10 quality point order #19 | Codex - GPT-5.6 Luna (high) Agent: Codex · Model variant: GPT-5.6 Luna (high) Agent version not exposed by source · source row 4bc4df53c3c0cc4b0ead07e1b4671e32 | 47.6 3-component source point | 53.4 | 71.8 | 17.5 | $0.96mean benchmark USD | 5.7 minmean agent wall time |
| #11 quality point order #35 | Claude Code - Kimi K2.6 Agent: Claude Code · Model variant: Kimi K2.6 Agent version not exposed by source · source row 1a928992a52f8b24fc9906536038249f | 30.6 3-component source point | 16.5 | 64.7 | 10.5 | $1.18mean benchmark USD | 41.2 minmean agent wall time |
| #12 quality point order #14 | Codex - GPT-5.6 Luna (xhigh) Agent: Codex · Model variant: GPT-5.6 Luna (xhigh) Agent version not exposed by source · source row f89183f8069a90af8146fa9d913d9f49 | 50.5 3-component source point | 56.6 | 76.2 | 18.8 | $1.26mean benchmark USD | 6.6 minmean agent wall time |
| #13 quality point order #26 | Codex - GPT-5.6 Sol (none) Agent: Codex · Model variant: GPT-5.6 Sol (none) Agent version not exposed by source · source row 1d09c3153c71c5f76dfe803461210bf9 | 39.5 3-component source point | 35.4 | 60.7 | 22.3 | $1.40mean benchmark USD | 3.4 minmean agent wall time |
| #14 quality point order #18 | Opencode - Muse Spark 1.1 (xhigh) Agent: Opencode · Model variant: Muse Spark 1.1 (xhigh) Agent version not exposed by source · source row 2de811eeba83a84002c01f63fe86c835 | 48.7 3-component source point | 54.3 | 73.0 | 18.8 | $1.43mean benchmark USD | 12.6 minmean agent wall time |
| #15 quality point order #11 | Codex - GPT-5.6 Luna (max) Agent: Codex · Model variant: GPT-5.6 Luna (max) Agent version not exposed by source · source row 76d8d6e68bbc152d859807f3f1d1fd05 | 53.8 3-component source point | 63.4 | 79.8 | 18.3 | $1.57mean benchmark USD | 8.0 minmean agent wall time |
| #16 quality point order #13 | Codex - GPT-5.6 Terra (high) Agent: Codex · Model variant: GPT-5.6 Terra (high) Agent version not exposed by source · source row 34efc84a2ae805e52ed2c70748ec2be7 | 51.3 3-component source point | 60.5 | 76.0 | 17.5 | $1.59mean benchmark USD | 6.2 minmean agent wall time |
| #17 quality point order #28 | Claude Code - Opus 4.7 (medium) Agent: Claude Code · Model variant: Opus 4.7 (medium) Agent version not exposed by source · source row 00c6d492a41655a179d9e181cddfa60e | 37.5 3-component source point | 27.4 | 71.4 | 13.7 | $1.68mean benchmark USD | 6.3 minmean agent wall time |
| #18 quality point order #12 | Codex - GPT-5.6 Terra (xhigh) Agent: Codex · Model variant: GPT-5.6 Terra (xhigh) Agent version not exposed by source · source row b8ac537bbb561ab3433a252c0f0fde8a | 52.7 3-component source point | 58.4 | 80.6 | 19.1 | $1.90mean benchmark USD | 6.9 minmean agent wall time |
| #19 quality point order #29 | Claude Code - Sonnet 4.6 (medium) Agent: Claude Code · Model variant: Sonnet 4.6 (medium) Agent version not exposed by source · source row 0d3005d0ab7dd58d68c092a2ef83f4b2 | 35.2 3-component source point | 28.9 | 63.5 | 13.2 | $1.97mean benchmark USD | 13.7 minmean agent wall time |
| #20 quality point order #36 | Gemini CLI - Gemini 3.1 Pro (high) Agent: Gemini CLI · Model variant: Gemini 3.1 Pro (high) Agent version not exposed by source · source row c3d8264417a3cd7cd792571f0795942d | 29.2 3-component source point | 14.2 | 68.3 | 5.1 | $2.00mean benchmark USD | 10.8 minmean agent wall time |
| #21 quality point order #23 | Cursor CLI - GPT-5.5 (medium) Agent: Cursor CLI · Model variant: GPT-5.5 (medium) Agent version not exposed by source · source row 5589375aeb2ef5e00064f40421d1112e | 42.8 3-component source point | 37.2 | 73.4 | 17.7 | $2.01mean benchmark USD | 6.6 minmean agent wall time |
| #22 quality point order #5 | Grok Build - Grok 4.5 (high) Agent: Grok Build · Model variant: Grok 4.5 (high) Agent version not exposed by source · source row 8d5e85efc646d9e29929f0519c2de4e0 | 57.9 3-component source point | 59.9 | 85.3 | 28.5 | $2.59mean benchmark USD | 16.5 minmean agent wall time |
| #23 quality point order #24 | Cursor CLI - Opus 4.7 (medium) Agent: Cursor CLI · Model variant: Opus 4.7 (medium) Agent version not exposed by source · source row 9dfbb045836418f93b25db6894649fc0 | 41.2 3-component source point | 31.6 | 70.6 | 21.5 | $2.68mean benchmark USD | 13.6 minmean agent wall time |
| #24 quality point order #15 | Codex - GPT-5.5 (medium) Agent: Codex · Model variant: GPT-5.5 (medium) Agent version not exposed by source · source row 8a576437d1fc29d09090c9523a1513f9 | 50.4 3-component source point | 56.6 | 75.8 | 18.8 | $2.75mean benchmark USD | 6.4 minmean agent wall time |
| #25 quality point order #6 | Codex - GPT-5.6 Terra (max) Agent: Codex · Model variant: GPT-5.6 Terra (max) Agent version not exposed by source · source row 9a837f89e0b12d31b1840f0cb88db3e9 | 57.4 3-component source point | 67.0 | 84.1 | 21.2 | $2.76mean benchmark USD | 8.4 minmean agent wall time |
| #26 quality point order #20 | Opencode - Opus 4.7 (medium) Agent: Opencode · Model variant: Opus 4.7 (medium) Agent version not exposed by source · source row 2c006e3b2d4b7047c65c655a9438b25e | 45.3 3-component source point | 39.5 | 75.0 | 21.5 | $2.93mean benchmark USD | 12.2 minmean agent wall time |
| #27 quality point order #7 | Kimi Code CLI - Kimi K3 Agent: Kimi Code CLI · Model variant: Kimi K3 Agent version not exposed by source · source row 95254776a26017018ef06a5c8fc90e47 | 56.8 3-component source point | 63.7 | 83.7 | 22.8 | $3.18mean benchmark USD | 23.8 minmean agent wall time |
| #28 quality point order #16 | Claude Code - Opus 4.8 (medium) Agent: Claude Code · Model variant: Opus 4.8 (medium) Agent version not exposed by source · source row 83e75dae6156204284804992d96ef400 | 48.9 3-component source point | 49.3 | 75.4 | 22.0 | $3.26mean benchmark USD | 12.4 minmean agent wall time |
| #29 quality point order #33 | Claude Code - GLM-5.1 Agent: Claude Code · Model variant: GLM-5.1 Agent version not exposed by source · source row 36fde43881babc825fc677be022de010 | 32.9 3-component source point | 18.6 | 65.1 | 15.1 | $4.33mean benchmark USD | 19.6 minmean agent wall time |
| #30 quality point order #8 | Codex - GPT-5.5 (xhigh) Agent: Codex · Model variant: GPT-5.5 (xhigh) Agent version not exposed by source · source row 822fe18f8eefb7fafe7014dcfb7f5be3 | 56.6 3-component source point | 64.3 | 84.1 | 21.5 | $5.07mean benchmark USD | 10.1 minmean agent wall time |
| #31 quality point order #21 | Claude Code - Opus 4.7 (max) Agent: Claude Code · Model variant: Opus 4.7 (max) Agent version not exposed by source · source row aadcd6ddf749fcd9a79721ae68081f20 | 45.2 3-component source point | 40.1 | 73.8 | 21.8 | $5.65mean benchmark USD | 15.8 minmean agent wall time |
| #32 quality point order #34 | Claude Code - Qwen3.7 Plus (thinking) Agent: Claude Code · Model variant: Qwen3.7 Plus (thinking) Agent version not exposed by source · source row ed32cad6a60710298313d063d29a961e | 32.5 3-component source point | 19.2 | 65.1 | 13.2 | $6.23mean benchmark USD | 10.6 minmean agent wall time |
| #33 quality point order #25 | Claude Code - GLM-5.2 Agent: Claude Code · Model variant: GLM-5.2 Agent version not exposed by source · source row 2270146beb4a66b898554e21e385c75d | 39.9 3-component source point | 28.6 | 72.3 | 18.8 | $6.47mean benchmark USD | 25.1 minmean agent wall time |
| #34 quality point order #1 | Codex - GPT-5.6 Sol (max) Agent: Codex · Model variant: GPT-5.6 Sol (max) Agent version not exposed by source · source row 6eb6667a6c986c2afc40c779a1666e5a | 61.0 3-component source point | 68.7 | 87.7 | 26.6 | $7.08mean benchmark USD | 10.2 minmean agent wall time |
| #35 quality point order #9 | Claude Code - Opus 4.8 (max) Agent: Claude Code · Model variant: Opus 4.8 (max) Agent version not exposed by source · source row 2b4a8c571a458c75b370cef5bcb37d17 | 54.9 3-component source point | 55.8 | 79.4 | 29.6 | $7.70mean benchmark USD | 23.1 minmean agent wall time |
| #36 quality point order #3 | Claude Code - Fable 5 (max) (with fallback) Agent: Claude Code · Model variant: Fable 5 (max) (with fallback) Agent version not exposed by source · source row 90c6cfde32ba917fece5bb788a5facd7 | 59.2 3-component source point | 66.1 | 82.5 | 29.0 | $11.72mean benchmark USD | 23.4 minmean agent wall time |
| — quality point order #2 | Codex - GPT-5.6 Sol (xhigh) Agent: Codex · Model variant: GPT-5.6 Sol (xhigh) Agent version not exposed by source · source row 5f0e4ab3a55661050fd03504cf608cb5 | 59.3 3-component source point | 67.0 | 86.1 | 24.7 | —not reported | 7.4 minmean agent wall time |
| — quality point order #4 | Codex - GPT-5.6 Sol (high) Agent: Codex · Model variant: GPT-5.6 Sol (high) Agent version not exposed by source · source row c9077abec519f68a3e8c8829f6a5bc4e | 58.1 3-component source point | 64.9 | 82.5 | 26.9 | —not reported | 6.3 minmean agent wall time |
| — quality point order #10 | Codex - GPT-5.6 Sol (medium) Agent: Codex · Model variant: GPT-5.6 Sol (medium) Agent version not exposed by source · source row f6d720f3864d8f1eeb50bf1905de9ed5 | 54.8 3-component source point | 64.0 | 77.8 | 22.6 | —not reported | 5.2 minmean agent wall time |
| — quality point order #17 | Codex - GPT-5.6 Sol (low) Agent: Codex · Model variant: GPT-5.6 Sol (low) Agent version not exposed by source · source row 33f9e014d631252d56034d11a88dc12a | 48.8 3-component source point | 53.4 | 73.0 | 19.9 | —not reported | 3.7 minmean agent wall time |
| — 2/3 legacy · unranked | Claude Code - Opus 4.6 (medium) Agent: Claude Code · Model variant: Opus 4.6 (medium) Agent version not exposed by source · source row b11f9d0125bd297e58bf1ae524918208 Retained legacy row has only 2/3 v1.2 components; DeepSWE is missing. Its displayed two-component headline is not compared with complete rows. | 42.4 2-component headline · unranked | Missing | 70.6 | 14.2 | $1.28mean benchmark USD | 8.0 minmean agent wall time |
| — 2/3 legacy · unranked | Codex - GPT-5.4 (medium) Agent: Codex · Model variant: GPT-5.4 (medium) Agent version not exposed by source · source row eef3c08a1f512ccfc36ff9c0307a600c Retained legacy row has only 2/3 v1.2 components; DeepSWE is missing. Its displayed two-component headline is not compared with complete rows. | 41.1 2-component headline · unranked | Missing | 69.8 | 12.4 | $2.27mean benchmark USD | 7.1 minmean agent wall time |
| — 2/3 legacy · unranked | Cursor CLI - Composer 2 Agent: Cursor CLI · Model variant: Composer 2 Agent version not exposed by source · source row 9d6039c9d6e56f451e737973c1bb11e9 Retained legacy row has only 2/3 v1.2 components; DeepSWE is missing. Its displayed two-component headline is not compared with complete rows. | 37.9 2-component headline · unranked | Missing | 64.7 | 11.0 | $0.04mean benchmark USD | 8.6 minmean agent wall time |
| — 2/3 legacy · unranked | Cursor CLI - GPT-5.4 (medium) Agent: Cursor CLI · Model variant: GPT-5.4 (medium) Agent version not exposed by source · source row 80c3216145b6845adf8d515966e0205d Retained legacy row has only 2/3 v1.2 components; DeepSWE is missing. Its displayed two-component headline is not compared with complete rows. | 40.1 2-component headline · unranked | Missing | 64.7 | 15.6 | $1.52mean benchmark USD | 8.3 minmean agent wall time |
All displayed fields come from the same Artificial Analysis Coding Agent Index v1.2 source snapshot. The four legacy rows appear after every complete row in all views and receive no quality, cost or speed position. Read the full source methodology →
What the score means
Simple average of pass@1 across DeepSWE, Terminal-Bench v2 and SWE-Atlas-QnA for rows with all three components. The source reports each benchmark component as pass@1 and its v1.2 index as their simple average.
These are point estimates only. We do not add an interval, significance label or rank-confidence range that the source does not provide.
Not part of the overall leaderboard
The source score evaluates complete coding-agent and model configurations on the Artificial Analysis v1.2 protocol. It is neither a bare-model score nor calibrated to SpringPrompt's cross-task predicted-fit scale.
Scale compatibility: not-proven
Frequently asked
Is this a ranking of the language models by themselves?
No. Every row preserves the complete source label: coding agent, model variant, reasoning qualifier and any fallback qualifier. Changing the agent can change the result.
Why are four rows unranked?
They are retained legacy rows with Terminal-Bench v2 and SWE-Atlas-QnA but no DeepSWE result. Their two-component headline is displayed for provenance but never compared with the complete three-component v1.2 cohort.
What do cheapest and fastest mean here?
They use Artificial Analysis's task-specific mean benchmark cost per task and mean agent wall time per task for the same source configuration. They do not change quality points.
Are score intervals or confidence ranks available?
No. The source exposes point estimates for this publication. SpringPrompt does not infer intervals, significance or rank confidence from those values.
Why is this excluded from the overall leaderboard?
The source score evaluates complete coding-agent and model configurations on the Artificial Analysis v1.2 protocol. It is neither a bare-model score nor calibrated to SpringPrompt's cross-task predicted-fit scale.