Provider and attribution
Artificial Analysis
- Registry state
- registered
- Rights state
- assumed
- Provider group
- artificial-analysis
Data sourced from Artificial Analysis
Release planning · no model scores
Exact V3 collection, facet, metric, lineage, identity and rights mapping for Coding and Coding Agent indices from Artificial Analysis.
Provider and attribution
Data sourced from Artificial Analysis
V3 evidence records
A repeated upstream benchmark can map to several collections. Each record below preserves its own task role, facets, capability atoms, subject identity, metrics, lineages and exclusions.
primary evidence · software-engineering/artificial-analysis/coding-and-coding-agent-indices
Which models are best at software engineering and coding?
Collection facets
Repository issue resolution · objective · 25%
Diagnose and fix real repository issues under an exact harness.
Implementation · hybrid · 25%
Build correct, scoped, maintainable changes.
Testing and debugging · hybrid · 25%
Find failures, write useful tests, and repair root causes.
Code review and security · hybrid · 25%
Spot correctness, security, and maintainability risks.
Capability atoms
Subject identity
Usable evidence lineages
deepswe/agentscientific-coding/foundation-modelswe-atlas-qna/agentterminal-bench/agentExact metric contract
| Metric | Lineage | Subject | Protocol | Use |
|---|---|---|---|---|
artificial_analysis_coding_indexArtificial Analysis Coding Index | aa-coding-aggregate | foundation-model | not supplied | aggregate-overlapaggregate overlaps selected component lineage(s): scientific-coding, terminal-bench |
scicodeSciCode | scientific-coding | foundation-model | SciCode under the Artificial Analysis evaluation protocol | usable |
deepsweDeepSWE | deepswe | agent | Artificial Analysis Coding Agent Index v1.1 DeepSWE; 113 tasks, 3 repeats, program-verifier pass@1 | usable |
terminal_bench_v2Terminal-Bench v2 (Coding Agent Index) | terminal-bench | agent | Artificial Analysis Coding Agent Index v1.1 Terminal-Bench v2; 84 compatible tasks, 3 repeats, test-suite pass@1 | usable |
swe_atlas_qnaSWE-Atlas-QnA | swe-atlas-qna | agent | Artificial Analysis Coding Agent Index v1.1 SWE-Atlas-QnA; 124 tasks, 3 repeats, rubric-scored pass@1 with partial credit | usable |
Exclusions and blockers
Mapping note
Do not count aggregate indices alongside their components.