Services · First packaged service · Available now
Catalog Content Quality
Evaluation and improvement of AI-generated catalog content inside PIM, product-feed, marketplace and ecommerce applications. We evaluate the software that transforms structured SKU information into titles, bullets, descriptions, channel listings and translations.
This is not another product-description generator. It is a controlled evaluation of the generation you already run, or the one you are about to buy, against your own source data and rules.
Why this workflow can be evaluated well
Structured product data gives us a source of truth.
Most generated business content is judged against taste. Catalog content is judged against a record: attributes, rules and channel constraints that already exist in your systems.
-
Claims can be checked
Every attribute in a listing either appears in the SKU record or it does not. Unsupported claims are detectable, not debatable.
-
Rules can be written down
Required and prohibited terms, brand and category rules and channel constraints are usually explicit already.
-
Generation repeats at scale
The same prompt runs across thousands of records, so a small error rate becomes a large publishing and review cost.
-
Teams already change models and prompts
Model releases, cost pressure and new channels force configuration changes that need a release decision.
-
Work can start offline
Representative records and rules are enough to begin; deep integration is not a prerequisite for the first answer.
-
The suite outlives the project
The reviewed test pack becomes a regression gate for every future prompt, model or provider change.
What we measure
Four families of criteria, deterministic checks first.
Schema, policy and source checks run before any model judgment. Subjective criteria are calibrated against your reviewers' accept, reject and edit decisions before they count.
Source fidelity
Deterministic + judged- Unsupported-claim rate
- Product-fact preservation
- Required-field coverage
Brand language
Deterministic + judged- Brand and terminology adherence
- Required and prohibited terms
- Category-specific conventions
Channel rules
Deterministic- Length and structure limits
- Banned phrasing and formatting
- Locale and marketplace constraints
Human-review burden
Observed- Publish-without-edit rate
- Edit distance
- Reviewer time per record
Every comparison also records cost and latency per publishable record, and reports performance separately on unseen categories, locales or channels. Critical failures such as invented product facts are reported on their own and block a recommendation regardless of the average score.
How an engagement runs
Define → Build → Compare → Prove
-
01
Define
Import representative product records, capture brand, category and channel rules, and write down the release decision, the scope, the critical failures and the threshold for a material improvement.
-
02
Build
Construct the evaluation set from real SKUs and consequential edge cases, review it with your team, and hold back cases for final validation.
-
03
Compare
Run your existing configuration and the alternatives on the same cases under equivalent conditions. Calibrate subjective criteria against human decisions and inspect failures one by one.
-
04
Prove
Test the recommended configuration on unseen data, then deliver the evidence packet and a reusable regression suite with the limitations stated.
What you provide
Inputs
- 01A representative sample of product records for the agreed categories, locales and channels
- 02Brand, category and channel rules, including required and prohibited terms
- 03Your current prompts, model and provider configuration, and sample outputs
- 04Examples your team has accepted, rejected and edited, where they exist
- 05The release decision you need to make and who makes it
What you receive
Deliverables
- 01An evaluation contract: population, slices, failure taxonomy, source of truth, criteria and critical failures
- 02A reviewed test pack built from your SKUs, with held-back cases for validation
- 03Baseline-versus-candidate results on the same cases, with cost and latency per publishable record
- 04A failure analysis that names specific failure modes rather than an average score
- 05A recommended configuration, validated on unseen data, with the decision to ship, reject, revise or investigate
- 06An evidence packet and a reusable regression suite your team can run on the next change
What this does not prove
The limits of offline evaluation, stated up front.
-
Publishability, not conversion
Offline quality against your rules and source data does not prove conversion, ranking or sales impact. Those require live measurement after release.
-
Your rules, your reviewers
Subjective criteria are calibrated to your team's decisions. A different brand or reviewer panel could reach a different threshold.
-
The agreed scope
Results apply to the categories, locales and channels in the evaluation set. Performance on unseen slices is reported separately and should be treated as provisional.
-
Text generation
The engagement evaluates text outputs from your generation configuration. It does not execute tools, retrieval pipelines or full agent workflows.
Next step
Tell us about the catalog, the channels and the decision you are stuck on.
We will reply with whether the engagement fits, what we would need from you, and how the scope would be agreed. Scope, duration and commercial terms are set per engagement.