Benchmarks / CatalogBench
CatalogBench
Which models can turn a sparse product feed, product photos and supplier copy into a listing that could go live, without inventing anything?
- Results dated
- 30 Sep 2026
- Models
- 18
- Unit
- % of missing fields
- Licence
- Spring Prompt original
- Judge
- openai/gpt-6.1-sol
- Runs
- 3 per model
| # | Model | CatalogBench: field accuracy % of missing fields, higher is better | Reliably publish-ready % of products | Publish-ready listings % of products | Reliably publish-ready % of products | Failed outputs % of products |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro PreviewGoogle |
95.4%
|
25.0% | 42.3% | 0.0% | 0.0% |
| 2 | Claude Fable 5.1Anthropic |
94.7%
|
16.1% | 35.1% | 0.0% | 0.0% |
| 3 | DeepSeek V4.1 FlashDeepSeek |
94.4%
|
35.7% | 58.3% | 10.7% | 0.0% |
| 3 | Kimi K3Moonshot AI |
94.4%
|
17.9% | 47.6% | 0.0% | 0.0% |
| 3 | GPT-6.1 SolOpenAI |
94.4%
|
69.6% | 74.4% | 66.1% | 0.0% |
| 6 | GPT-6 AstraOpenAI |
94.2%
|
71.4% | 74.4% | 66.1% | 0.0% |
| 7 | Claude Haiku 4.5Anthropic |
93.7%
|
0.0% | 0.6% | 0.0% | 0.0% |
| 7 | GPT-6 LunaOpenAI |
93.7%
|
51.8% | 64.3% | 42.9% | 0.0% |
| 7 | GPT-6 SolOpenAI |
93.7%
|
60.7% | 70.2% | 55.4% | 0.0% |
| 10 | Muse Spark 1.3Meta |
92.9%
|
37.5% | 57.7% | 7.1% | 0.0% |
| 10 | Grok 4.7xAI |
92.9%
|
44.6% | 62.5% | 10.7% | 0.0% |
| 12 | Gemini 3.8 FlashGoogle |
92.7%
|
16.1% | 33.3% | 1.8% | 0.0% |
| 13 | Gemini 3.5 Flash LiteGoogle |
92.5%
|
3.6% | 15.5% | 0.0% | 0.0% |
| 14 | Claude Opus 5.5Anthropic |
92.2%
|
39.3% | 58.3% | 0.0% | 0.0% |
| 15 | Mistral Medium 3.5Mistral |
91.7%
|
5.4% | 15.5% | 0.0% | 0.0% |
| 16 | Claude Sonnet 5.5Anthropic |
91.5%
|
21.4% | 35.1% | 1.8% | 0.0% |
| 17 | Qwen3.8 Max (0902)AlibabaFailed outputs |
90.3%
|
10.7% | 28.0% | 0.0% | 3.6% |
| 18 | GLM 5V TurboZ.aiFailed outputs |
59.6%
|
0.0% | 3.0% | 0.0% | 32.1% |
Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.
What it measures
- Attributes read from the images, not guessed
- Supplier claims checked, not repeated
- Required UK product information included
- Listings that shoppers can find in search
What it does not measure
- Conversion or sales impact
- Real product photography (a real-photo slice is planned)
- Writing style beyond the listed checks
Method
- Rule-based checks first; judged checks are yes or no
- The judge was checked for bias against Gemini and Claude judges
- Private products are held back so the set can be refreshed
Checking the judge
The judge is an OpenAI model, and OpenAI models lead this table, so we checked it for bias. Gemini 3.1 Pro and Claude Opus 5.5 judged the same outputs from five models on a calibration set. All three judges put the models in the same order under both briefs. Each was slightly gentler on its own family's marketing copy: the top GPT models moved by 4 to 8 points between judges under the marketing brief, without changing places.
Failures
Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.
- Qwen3.8 Max (0902): reply cut off at the token limit: 6 of 168 attempts.
- GLM 5V Turbo: invalid JSON (raw line breaks inside text): 50 of 168 attempts; invalid JSON: 4 of 168 attempts.