Confirm Action

Are you sure you want to proceed?

Attributed direct benchmark · agent configurations

Best coding agent configurations

Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.

Codex - GPT-5.6 Sol (max) leads the complete v1.2 point order at 61.0.

Rows are source-listed coding-agent and model configurations. Artificial Analysis does not expose agent harness versions, so these are not bare-model ranks or immutable runnable-system identities. Quality, cost and wall time are publisher point estimates; no interval, significance or rank-confidence claim is available.

Source configurations

44

Complete 3/3 cohort

40

Legacy 2/3 · unranked

4

Lowest source-reported benchmark cost per task

Positions use the publisher's mean benchmark cost per task for the same exact agent configuration. Cost changes no quality score. Complete rows without a cost remain visible but unpositioned.

Source snapshot 2026-07-20

Lowest source-reported benchmark cost per task Exact agent + model configuration Index point score DeepSWE Terminal-Bench v2 SWE-Atlas-QnA Cost/task Wall time/task
#1 quality point order #31 Cursor CLI - Composer 2.5 Agent: Cursor CLI · Model variant: Composer 2.5 Agent version not exposed by source · source row 7e4d99f22241d821b64798b319e655de 33.8 3-component source point 15.9 67.1 18.3 $0.08mean benchmark USD 9.6 minmean agent wall time
#2 quality point order #38 Codex - GPT-5.6 Luna (low) Agent: Codex · Model variant: GPT-5.6 Luna (low) Agent version not exposed by source · source row ba14c764e7104dbf50b2f160ebf14426 23.4 3-component source point 10.3 49.6 10.2 $0.21mean benchmark USD 1.9 minmean agent wall time
#3 quality point order #37 Claude Code - DeepSeek V4 Pro (high) Agent: Claude Code · Model variant: DeepSeek V4 Pro (high) Agent version not exposed by source · source row 19819b29f8069ec81aec31ce3fa68958 28.5 3-component source point 8.6 65.9 11.0 $0.27mean benchmark USD 17.9 minmean agent wall time
#4 quality point order #40 Codex - GPT-5.6 Luna (none) Agent: Codex · Model variant: GPT-5.6 Luna (none) Agent version not exposed by source · source row 824070d2da3d80c2664a07004bf77f36 17.9 3-component source point 6.5 37.3 9.9 $0.35mean benchmark USD 2.5 minmean agent wall time
#5 quality point order #39 Codex - GPT-5.6 Terra (none) Agent: Codex · Model variant: GPT-5.6 Terra (none) Agent version not exposed by source · source row 0a3201d26f6411e9328800c6bfbcc1e2 21.9 3-component source point 13.3 39.3 13.2 $0.37mean benchmark USD 1.8 minmean agent wall time
#6 quality point order #27 Codex - GPT-5.6 Luna (medium) Agent: Codex · Model variant: GPT-5.6 Luna (medium) Agent version not exposed by source · source row 8c17551948c19f8277b123a445a01485 38.4 3-component source point 36.6 63.5 15.1 $0.47mean benchmark USD 3.4 minmean agent wall time
#7 quality point order #30 Codex - GPT-5.6 Terra (low) Agent: Codex · Model variant: GPT-5.6 Terra (low) Agent version not exposed by source · source row deffe38fe77a5fd77c871140402adf8f 33.9 3-component source point 29.8 57.5 14.2 $0.48mean benchmark USD 2.8 minmean agent wall time
#8 quality point order #32 Cursor CLI - Composer 2.5 Fast Agent: Cursor CLI · Model variant: Composer 2.5 Fast Agent version not exposed by source · source row 913859e7107291a71a76e2ad0bac704a 33.8 3-component source point 15.9 67.1 18.3 $0.56mean benchmark USD 6.8 minmean agent wall time
#9 quality point order #22 Codex - GPT-5.6 Terra (medium) Agent: Codex · Model variant: GPT-5.6 Terra (medium) Agent version not exposed by source · source row 9cd9393327109ca9132fff0834a65880 43.9 3-component source point 45.7 69.4 16.7 $0.90mean benchmark USD 4.3 minmean agent wall time
#10 quality point order #19 Codex - GPT-5.6 Luna (high) Agent: Codex · Model variant: GPT-5.6 Luna (high) Agent version not exposed by source · source row 4bc4df53c3c0cc4b0ead07e1b4671e32 47.6 3-component source point 53.4 71.8 17.5 $0.96mean benchmark USD 5.7 minmean agent wall time
#11 quality point order #35 Claude Code - Kimi K2.6 Agent: Claude Code · Model variant: Kimi K2.6 Agent version not exposed by source · source row 1a928992a52f8b24fc9906536038249f 30.6 3-component source point 16.5 64.7 10.5 $1.18mean benchmark USD 41.2 minmean agent wall time
#12 quality point order #14 Codex - GPT-5.6 Luna (xhigh) Agent: Codex · Model variant: GPT-5.6 Luna (xhigh) Agent version not exposed by source · source row f89183f8069a90af8146fa9d913d9f49 50.5 3-component source point 56.6 76.2 18.8 $1.26mean benchmark USD 6.6 minmean agent wall time
#13 quality point order #26 Codex - GPT-5.6 Sol (none) Agent: Codex · Model variant: GPT-5.6 Sol (none) Agent version not exposed by source · source row 1d09c3153c71c5f76dfe803461210bf9 39.5 3-component source point 35.4 60.7 22.3 $1.40mean benchmark USD 3.4 minmean agent wall time
#14 quality point order #18 Opencode - Muse Spark 1.1 (xhigh) Agent: Opencode · Model variant: Muse Spark 1.1 (xhigh) Agent version not exposed by source · source row 2de811eeba83a84002c01f63fe86c835 48.7 3-component source point 54.3 73.0 18.8 $1.43mean benchmark USD 12.6 minmean agent wall time
#15 quality point order #11 Codex - GPT-5.6 Luna (max) Agent: Codex · Model variant: GPT-5.6 Luna (max) Agent version not exposed by source · source row 76d8d6e68bbc152d859807f3f1d1fd05 53.8 3-component source point 63.4 79.8 18.3 $1.57mean benchmark USD 8.0 minmean agent wall time
#16 quality point order #13 Codex - GPT-5.6 Terra (high) Agent: Codex · Model variant: GPT-5.6 Terra (high) Agent version not exposed by source · source row 34efc84a2ae805e52ed2c70748ec2be7 51.3 3-component source point 60.5 76.0 17.5 $1.59mean benchmark USD 6.2 minmean agent wall time
#17 quality point order #28 Claude Code - Opus 4.7 (medium) Agent: Claude Code · Model variant: Opus 4.7 (medium) Agent version not exposed by source · source row 00c6d492a41655a179d9e181cddfa60e 37.5 3-component source point 27.4 71.4 13.7 $1.68mean benchmark USD 6.3 minmean agent wall time
#18 quality point order #12 Codex - GPT-5.6 Terra (xhigh) Agent: Codex · Model variant: GPT-5.6 Terra (xhigh) Agent version not exposed by source · source row b8ac537bbb561ab3433a252c0f0fde8a 52.7 3-component source point 58.4 80.6 19.1 $1.90mean benchmark USD 6.9 minmean agent wall time
#19 quality point order #29 Claude Code - Sonnet 4.6 (medium) Agent: Claude Code · Model variant: Sonnet 4.6 (medium) Agent version not exposed by source · source row 0d3005d0ab7dd58d68c092a2ef83f4b2 35.2 3-component source point 28.9 63.5 13.2 $1.97mean benchmark USD 13.7 minmean agent wall time
#20 quality point order #36 Gemini CLI - Gemini 3.1 Pro (high) Agent: Gemini CLI · Model variant: Gemini 3.1 Pro (high) Agent version not exposed by source · source row c3d8264417a3cd7cd792571f0795942d 29.2 3-component source point 14.2 68.3 5.1 $2.00mean benchmark USD 10.8 minmean agent wall time
#21 quality point order #23 Cursor CLI - GPT-5.5 (medium) Agent: Cursor CLI · Model variant: GPT-5.5 (medium) Agent version not exposed by source · source row 5589375aeb2ef5e00064f40421d1112e 42.8 3-component source point 37.2 73.4 17.7 $2.01mean benchmark USD 6.6 minmean agent wall time
#22 quality point order #5 Grok Build - Grok 4.5 (high) Agent: Grok Build · Model variant: Grok 4.5 (high) Agent version not exposed by source · source row 8d5e85efc646d9e29929f0519c2de4e0 57.9 3-component source point 59.9 85.3 28.5 $2.59mean benchmark USD 16.5 minmean agent wall time
#23 quality point order #24 Cursor CLI - Opus 4.7 (medium) Agent: Cursor CLI · Model variant: Opus 4.7 (medium) Agent version not exposed by source · source row 9dfbb045836418f93b25db6894649fc0 41.2 3-component source point 31.6 70.6 21.5 $2.68mean benchmark USD 13.6 minmean agent wall time
#24 quality point order #15 Codex - GPT-5.5 (medium) Agent: Codex · Model variant: GPT-5.5 (medium) Agent version not exposed by source · source row 8a576437d1fc29d09090c9523a1513f9 50.4 3-component source point 56.6 75.8 18.8 $2.75mean benchmark USD 6.4 minmean agent wall time
#25 quality point order #6 Codex - GPT-5.6 Terra (max) Agent: Codex · Model variant: GPT-5.6 Terra (max) Agent version not exposed by source · source row 9a837f89e0b12d31b1840f0cb88db3e9 57.4 3-component source point 67.0 84.1 21.2 $2.76mean benchmark USD 8.4 minmean agent wall time
#26 quality point order #20 Opencode - Opus 4.7 (medium) Agent: Opencode · Model variant: Opus 4.7 (medium) Agent version not exposed by source · source row 2c006e3b2d4b7047c65c655a9438b25e 45.3 3-component source point 39.5 75.0 21.5 $2.93mean benchmark USD 12.2 minmean agent wall time
#27 quality point order #7 Kimi Code CLI - Kimi K3 Agent: Kimi Code CLI · Model variant: Kimi K3 Agent version not exposed by source · source row 95254776a26017018ef06a5c8fc90e47 56.8 3-component source point 63.7 83.7 22.8 $3.18mean benchmark USD 23.8 minmean agent wall time
#28 quality point order #16 Claude Code - Opus 4.8 (medium) Agent: Claude Code · Model variant: Opus 4.8 (medium) Agent version not exposed by source · source row 83e75dae6156204284804992d96ef400 48.9 3-component source point 49.3 75.4 22.0 $3.26mean benchmark USD 12.4 minmean agent wall time
#29 quality point order #33 Claude Code - GLM-5.1 Agent: Claude Code · Model variant: GLM-5.1 Agent version not exposed by source · source row 36fde43881babc825fc677be022de010 32.9 3-component source point 18.6 65.1 15.1 $4.33mean benchmark USD 19.6 minmean agent wall time
#30 quality point order #8 Codex - GPT-5.5 (xhigh) Agent: Codex · Model variant: GPT-5.5 (xhigh) Agent version not exposed by source · source row 822fe18f8eefb7fafe7014dcfb7f5be3 56.6 3-component source point 64.3 84.1 21.5 $5.07mean benchmark USD 10.1 minmean agent wall time
#31 quality point order #21 Claude Code - Opus 4.7 (max) Agent: Claude Code · Model variant: Opus 4.7 (max) Agent version not exposed by source · source row aadcd6ddf749fcd9a79721ae68081f20 45.2 3-component source point 40.1 73.8 21.8 $5.65mean benchmark USD 15.8 minmean agent wall time
#32 quality point order #34 Claude Code - Qwen3.7 Plus (thinking) Agent: Claude Code · Model variant: Qwen3.7 Plus (thinking) Agent version not exposed by source · source row ed32cad6a60710298313d063d29a961e 32.5 3-component source point 19.2 65.1 13.2 $6.23mean benchmark USD 10.6 minmean agent wall time
#33 quality point order #25 Claude Code - GLM-5.2 Agent: Claude Code · Model variant: GLM-5.2 Agent version not exposed by source · source row 2270146beb4a66b898554e21e385c75d 39.9 3-component source point 28.6 72.3 18.8 $6.47mean benchmark USD 25.1 minmean agent wall time
#34 quality point order #1 Codex - GPT-5.6 Sol (max) Agent: Codex · Model variant: GPT-5.6 Sol (max) Agent version not exposed by source · source row 6eb6667a6c986c2afc40c779a1666e5a 61.0 3-component source point 68.7 87.7 26.6 $7.08mean benchmark USD 10.2 minmean agent wall time
#35 quality point order #9 Claude Code - Opus 4.8 (max) Agent: Claude Code · Model variant: Opus 4.8 (max) Agent version not exposed by source · source row 2b4a8c571a458c75b370cef5bcb37d17 54.9 3-component source point 55.8 79.4 29.6 $7.70mean benchmark USD 23.1 minmean agent wall time
#36 quality point order #3 Claude Code - Fable 5 (max) (with fallback) Agent: Claude Code · Model variant: Fable 5 (max) (with fallback) Agent version not exposed by source · source row 90c6cfde32ba917fece5bb788a5facd7 59.2 3-component source point 66.1 82.5 29.0 $11.72mean benchmark USD 23.4 minmean agent wall time
quality point order #2 Codex - GPT-5.6 Sol (xhigh) Agent: Codex · Model variant: GPT-5.6 Sol (xhigh) Agent version not exposed by source · source row 5f0e4ab3a55661050fd03504cf608cb5 59.3 3-component source point 67.0 86.1 24.7 not reported 7.4 minmean agent wall time
quality point order #4 Codex - GPT-5.6 Sol (high) Agent: Codex · Model variant: GPT-5.6 Sol (high) Agent version not exposed by source · source row c9077abec519f68a3e8c8829f6a5bc4e 58.1 3-component source point 64.9 82.5 26.9 not reported 6.3 minmean agent wall time
quality point order #10 Codex - GPT-5.6 Sol (medium) Agent: Codex · Model variant: GPT-5.6 Sol (medium) Agent version not exposed by source · source row f6d720f3864d8f1eeb50bf1905de9ed5 54.8 3-component source point 64.0 77.8 22.6 not reported 5.2 minmean agent wall time
quality point order #17 Codex - GPT-5.6 Sol (low) Agent: Codex · Model variant: GPT-5.6 Sol (low) Agent version not exposed by source · source row 33f9e014d631252d56034d11a88dc12a 48.8 3-component source point 53.4 73.0 19.9 not reported 3.7 minmean agent wall time
2/3 legacy · unranked Claude Code - Opus 4.6 (medium) Agent: Claude Code · Model variant: Opus 4.6 (medium) Agent version not exposed by source · source row b11f9d0125bd297e58bf1ae524918208 Retained legacy row has only 2/3 v1.2 components; DeepSWE is missing. Its displayed two-component headline is not compared with complete rows. 42.4 2-component headline · unranked Missing 70.6 14.2 $1.28mean benchmark USD 8.0 minmean agent wall time
2/3 legacy · unranked Codex - GPT-5.4 (medium) Agent: Codex · Model variant: GPT-5.4 (medium) Agent version not exposed by source · source row eef3c08a1f512ccfc36ff9c0307a600c Retained legacy row has only 2/3 v1.2 components; DeepSWE is missing. Its displayed two-component headline is not compared with complete rows. 41.1 2-component headline · unranked Missing 69.8 12.4 $2.27mean benchmark USD 7.1 minmean agent wall time
2/3 legacy · unranked Cursor CLI - Composer 2 Agent: Cursor CLI · Model variant: Composer 2 Agent version not exposed by source · source row 9d6039c9d6e56f451e737973c1bb11e9 Retained legacy row has only 2/3 v1.2 components; DeepSWE is missing. Its displayed two-component headline is not compared with complete rows. 37.9 2-component headline · unranked Missing 64.7 11.0 $0.04mean benchmark USD 8.6 minmean agent wall time
2/3 legacy · unranked Cursor CLI - GPT-5.4 (medium) Agent: Cursor CLI · Model variant: GPT-5.4 (medium) Agent version not exposed by source · source row 80c3216145b6845adf8d515966e0205d Retained legacy row has only 2/3 v1.2 components; DeepSWE is missing. Its displayed two-component headline is not compared with complete rows. 40.1 2-component headline · unranked Missing 64.7 15.6 $1.52mean benchmark USD 8.3 minmean agent wall time

All displayed fields come from the same Artificial Analysis Coding Agent Index v1.2 source snapshot. The four legacy rows appear after every complete row in all views and receive no quality, cost or speed position. Read the full source methodology →

What the score means

Simple average of pass@1 across DeepSWE, Terminal-Bench v2 and SWE-Atlas-QnA for rows with all three components. The source reports each benchmark component as pass@1 and its v1.2 index as their simple average.

These are point estimates only. We do not add an interval, significance label or rank-confidence range that the source does not provide.

Not part of the overall leaderboard

The source score evaluates complete coding-agent and model configurations on the Artificial Analysis v1.2 protocol. It is neither a bare-model score nor calibrated to SpringPrompt's cross-task predicted-fit scale.

Scale compatibility: not-proven

Frequently asked

Is this a ranking of the language models by themselves?

No. Every row preserves the complete source label: coding agent, model variant, reasoning qualifier and any fallback qualifier. Changing the agent can change the result.

Why are four rows unranked?

They are retained legacy rows with Terminal-Bench v2 and SWE-Atlas-QnA but no DeepSWE result. Their two-component headline is displayed for provenance but never compared with the complete three-component v1.2 cohort.

What do cheapest and fastest mean here?

They use Artificial Analysis's task-specific mean benchmark cost per task and mean agent wall time per task for the same source configuration. They do not change quality points.

Are score intervals or confidence ranks available?

No. The source exposes point estimates for this publication. SpringPrompt does not infer intervals, significance or rank confidence from those values.

Why is this excluded from the overall leaderboard?

The source score evaluates complete coding-agent and model configurations on the Artificial Analysis v1.2 protocol. It is neither a bare-model score nor calibrated to SpringPrompt's cross-task predicted-fit scale.