What this benchmark tests
A legacy terminal-use evaluation of agentic coding, system administration, data processing, and related command-line tasks. Artificial Analysis marks this track as superseded by Terminal-Bench v2.1; historical scores remain useful only within the legacy setup.
How to read the score
The published metric is Pass rate. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.
A missing source value is not scored as zero: that configuration is omitted from this benchmark page.
Comparability policy
Why this is a system evaluation
A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.
That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.
What can be compared here
Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same terminalbenchHard field.
This filtered table contains only the published Nemotron 3 Ultra 550B A55B (Reasoning) configurations from the complete official comparison. It does not hide other model families from the benchmark page or merge results from a different evaluation track.