How we measure, and how to read the results.
Every number on this site says what was measured, by whom, on what, and when, and what it does not show.
Our principles
- Start with the decision, not the metric. A benchmark is only useful if it helps someone choose.
- Rule-based checks first. Where a rule can decide, it does. AI judges are used only where rules cannot decide, and are checked against people's decisions.
- Check the checker. Two judges agreeing shows they are consistent, not that they are right.
- Report failures, not just averages. A critical failure is shown even when the average looks good.
- Test the whole system. Prompt, model, settings and provider together decide the result.
- Keep uncertainty visible. Ranges and unresolved cases are shown, not smoothed away.
- Keep the failures. Every confirmed failure becomes a permanent test.
- Count the cost of a decision, not just the cost of a model call.
Measured by us, or reported by others
Results from our own benchmarks are labelled Measured by Spring Prompt. We ran them, and we keep every model output so each result can be traced back to what the model actually wrote.
Results from other benchmarks are labelled Reported by the source, with its credit and licence. We do not re-run them, and we only publish sources whose licence allows it.
We never blend sources into a single overall score. Different benchmarks measure different things, in different units.
Reading a results table
- Numbers are shown in the source's own units, with its date.
- Where a source gives a range (for example a 95% interval), models whose ranges overlap share a rank. A rank is not a verdict that one model is better.
- Each model runs at its provider's default settings unless the table says otherwise.
- A result says nothing about work the benchmark does not test. Each benchmark page lists what it does not measure.
Our own benchmarks
CatalogBench tests product-listing enrichment from feeds, images and supplier copy. ROASBench tests a year of marketing budget decisions in a simulated market. BulletBench tests decisions under a real clock. Each page explains its method and limits.
Where an AI judges part of a result, we name the judge and check it for bias. For CatalogBench we re-scored a calibration run with judges from Google and Anthropic, and the order of the models did not change.
What gets published, and indexed
Pages are built from sealed, versioned releases of our catalogue. A page is only offered to search engines when it has enough evidence to be useful: for example, a comparison needs both models on at least three shared sources. Thinner pages still exist, but are not indexed.
Corrections
If you find an error, tell us. We correct it in the next release and keep the previous release, so every change can be traced.