Confirm Action

Are you sure you want to proceed?

Release planning · no model scores

Coding and Coding Agent indices evidence mapping

Exact V3 collection, facet, metric, lineage, identity and rights mapping for Coding and Coding Agent indices from Artificial Analysis.

V3 evidence records

1 collection mapping

A repeated upstream benchmark can map to several collections. Each record below preserves its own task role, facets, capability atoms, subject identity, metrics, lineages and exclusions.

primary evidence · software-engineering/artificial-analysis/coding-and-coding-agent-indices

Software Engineering & Coding

Which models are best at software engineering and coding?

Upstream access
machine-ready
Eligibility
runnable-now
Registry
registered
Max task directness
70%

Collection facets

  • Repository issue resolution · objective · 25%

    Diagnose and fix real repository issues under an exact harness.

  • Implementation · hybrid · 25%

    Build correct, scoped, maintainable changes.

  • Testing and debugging · hybrid · 25%

    Find failures, write useful tests, and repair root causes.

  • Code review and security · hybrid · 25%

    Spot correctness, security, and maintainability risks.

Capability atoms

codingrepository_engineeringterminal_workrepository_understanding

Subject identity

Declared:
harnessed-model
Metric partitions:
agent, foundation-model
Separate partition required:
yes
Identity projection:
forbidden

Usable evidence lineages

  • deepswe/agent
  • scientific-coding/foundation-model
  • swe-atlas-qna/agent
  • terminal-bench/agent

Exact metric contract

MetricLineageSubjectProtocolUse
artificial_analysis_coding_indexArtificial Analysis Coding Indexaa-coding-aggregatefoundation-modelnot suppliedaggregate-overlapaggregate overlaps selected component lineage(s): scientific-coding, terminal-bench
scicodeSciCodescientific-codingfoundation-modelSciCode under the Artificial Analysis evaluation protocolusable
deepsweDeepSWEdeepsweagentArtificial Analysis Coding Agent Index v1.1 DeepSWE; 113 tasks, 3 repeats, program-verifier pass@1usable
terminal_bench_v2Terminal-Bench v2 (Coding Agent Index)terminal-benchagentArtificial Analysis Coding Agent Index v1.1 Terminal-Bench v2; 84 compatible tasks, 3 repeats, test-suite pass@1usable
swe_atlas_qnaSWE-Atlas-QnAswe-atlas-qnaagentArtificial Analysis Coding Agent Index v1.1 SWE-Atlas-QnA; 124 tasks, 3 repeats, rubric-scored pass@1 with partial creditusable

Exclusions and blockers

  • source redistribution is conditional: assumed
  • metric subject identity differs from the editorial mapping; publish and calibrate each identity partition separately
  • artificial_analysis_coding_index: aggregate overlaps selected component lineage(s): scientific-coding, terminal-bench

Mapping note

Do not count aggregate indices alongside their components.