AI COMMERCE QA FOR ECOMMERCE AGENCIES

Test whether your clients' product pages can support an AI shopping task.

AisleLens runs defined buyer tests against real storefront evidence. See what the store can answer, where the evidence runs out, what your agency can correct, and whether the same test passes after the change.

One product. One buying task. Every requirement proven, not proven, or requires store access — with the sentence that decided it.

Run a real test · See a complete example → · Read Coffee Standard v1.3 → · All published standards · Methodology

What your agency hands the client.

Six artifacts, all of them checkable by someone who was not in the room.

A client-ready evidence audit
Every requirement in the standard, with the store's own sentence beside it and the surface that sentence was read from. Nothing to take on trust.
The exact buyer questions the store cannot answer
Not a category of weakness — the specific questions, quoted from a published standard, that public evidence could not settle on this page.
A line between what the store controls and what it does not
Some failures are a missing sentence on a product page. Some are outside the store entirely. The report says which, and refuses to propose an edit it cannot justify.
Corrections tied to the assertion that failed
Each proposed change names the requirement it is meant to satisfy and the evidence form the standard accepts for it. Reviewed and approved by you, and reversible.
A before-and-after rerun
The identical test, same standard, same content hash, run again after the change — reported either way, including when nothing moved.
A regression baseline you keep
A requirement that passes becomes a check that keeps running, so a theme update or a catalog edit that undoes the work is visible rather than silent.

Test · Trace · Correct · Rerun · Retain

The same five steps on every client, in the same order, with an artifact at each one.

  1. Test A published buying standard for the category is executed against the client's live product page.
  2. Trace Every result names the surface it was read from and quotes the sentence that decided it, or states that no sentence existed.
  3. Correct The failures the store controls get a proposed, reversible change, tied to the requirement it is meant to satisfy.
  4. Rerun The identical test runs again against the same version and content hash, so the question cannot have moved between the two runs.
  5. Retain The passing test stays as a regression check, and reports when a later change takes it back.

What an executable buyer test actually does.

A buying task becomes a list of assertions. Each assertion is settled from evidence that is retrieved, quoted and attributed — or it is not settled, and the result says so.

Product description
The readable copy a shopper sees, sentence by sentence.
Options and variants
The purchasable option list, and whether the matching variant is actually available.
Structured data
The JSON-LD product node — identifiers, offers, availability, category.
Policy pages
Shipping and returns, fetched separately and attributed separately.
Page metadata
Title, canonical, description — the machine-facing summary of the page.
Authorized store data
Only with a connected store, and only where public data provably cannot settle the question.

And where it stops. The evaluator is deterministic: it matches evidence, it does not reason about the product. A requirement with no retrievable sentence behind it is reported as not proven, never inferred from context, never softened into a maybe. A requirement that public data cannot settle at all is reported as requires store access — a third state that exists so the first two stay honest.

One real store, every row, nothing selected for effect.

This is the complete result the published standard produces on a real coffee product page — not an excerpt chosen to flatter either side. Each row cites the entry it executes, at a version and a content hash that still resolve.

A REAL RESULT, ON A REAL STORE: Klatch Coffee — Ethiopia Yirgacheffe Supernatural (https://klatchcoffee.com/products/ethiopia-yirgacheffe-supernatural)

Standard: AisleLens Coffee Standard v1.3 · content hash ba2050578ed0274885fd6213967c230b2a57dd2b7c1d3fba8c5e1633027d4cf7 · captured 2026-07-27T04:53:11.912Z

5 proven · 5 not proven · 0 requires store access · 10 requirements asked

Replayed offline from a frozen capture of the live page, against the published standard named above. Every row links to the entry it executes.

✓ proven · ✕ not proven · – no blocking evidence · ○ requires store access

  1. ✓ Can I buy this as whole beans?

    A "Whole Bean" variant is listed and purchasable.

    73 of 99 coffee stores don't state this. This one does.

    Entry: ALS-COFFEE-1.3-FORMAT-001

  2. ✕ Can I buy this already ground, so I do not need a grinder?

    Checked the public variant list, and no "Ground" variant is offered on this product.

    92 of 99 coffee stores don't state this either.

    Entry: ALS-COFFEE-1.3-FORMAT-002

  3. ✓ Can I get this ground for espresso?

    A "Espresso" variant is listed and purchasable.

    84 of 99 coffee stores don't state this. This one does.

    Entry: ALS-COFFEE-1.3-GRIND-001

  4. ✓ Can I get this ground for a filter or pour-over brewer?

    A "Filter" variant is listed and purchasable.

    92 of 99 coffee stores don't state this. This one does.

    Entry: ALS-COFFEE-1.3-GRIND-002

  5. ✓ How much coffee do I get — is there a weight anywhere on this page?

    Stated in your variant options.

    2 lb bag.

    Read from: variant options — https://klatchcoffee.com/products/ethiopia-yirgacheffe-supernatural

    49 of 100 coffee stores don't state this. This one does.

    Entry: ALS-COFFEE-1.3-WEIGHT-001

  6. ✕ Does this page say the coffee is organic?

    Checked product copy, structured data, product title, variant options and page description, and none of them state an organic claim in wording an AI buyer could verify.

    92 of 100 coffee stores don't state this either.

    Entry: ALS-COFFEE-1.3-CERT-001

  7. ✕ Does this page say the coffee is fair trade?

    Checked product copy, structured data, product title, variant options and page description, and none of them state a fair trade claim in wording an AI buyer could verify.

    96 of 100 coffee stores don't state this either.

    Entry: ALS-COFFEE-1.3-CERT-002

  8. ✕ Is this coffee from one place, or is it a blend?

    Checked product copy, structured data, product title, variant options and page description, and none of them state a single origin claim in wording an AI buyer could verify.

    89 of 100 coffee stores don't state this either.

    Entry: ALS-COFFEE-1.3-SOURCE-001

  9. ✕ Can a shopping assistant match this exact bag to a catalogue entry?

    Your product structured data publishes no GTIN or MPN, so a machine buyer can't match this product to a catalogue entry.

    74 of 76 coffee stores don't state this either.

    Entry: ALS-COFFEE-1.3-IDENT-001

  10. ✓ When will this actually be sent to me?

    Delivery timing is stated in your shipping policy.

    Domestic Shipping Policy Shipment processing time All orders are processed within 2-3 business days - please keep in mind that we do fresh-roast all coffee prior to shipment after…

    Read from: shipping policy — https://www.klatchcoffee.com/policies/shipping-policy

    45 of the 74 coffee stores we could decide (of 100 asked) don't state this. This one does.

    Entry: ALS-COFFEE-1.3-DELIV-001

Where the standard has published a measurement for an entry, the row says how the rest of the sample did on the same question — with the number of stores that question could actually be decided on, which is not always the whole sample.

Read the complete test →

The result moves when the evidence moves. Never when the question does.

Same standard. Same version. Same content hash. Same entry id. The only difference between the two columns is a sentence on the product page.

Unchanged across the rerun: AisleLens Coffee Standard v1.3, content hash ba2050578ed0274885fd6213967c230b2a57dd2b7c1d3fba8c5e1633027d4cf7, entry ALS-COFFEE-1.3-FORMAT-002, question “Can I buy this already ground, so I do not need a grinder?”.

Before — the page as it is today: ✕ not proven — Checked the public variant list, and no "Ground" variant is offered on this product.

After — the page with the evidence the standard accepts: ✓ proven — the page states a purchasable option value naming the ground state, for example “Ground”.

The right-hand sentence is illustrative: it is the accepted-evidence example the standard itself publishes for this entry, not text from any store. The left-hand column is the real current result.

One failed test. One isolated cause. One verified rerun.

  • Before the fix: 0 of 4 test runs passed — the required claim could not be verified from any store surface.
  • After one approved, reversible correction: 4 of 4 passed. Same test, same models, versions pinned.
  • Unsupported evidence credited: zero — every claim in every run traces to retrieved evidence.

This is a controlled technical validation on a Shopify development store, labeled as such. It is not a merchant case, and nothing on this page presents it as one.

A summary number tells you that something moved. A test tells you what broke.

Mention monitoringReadiness checklistsAisleLens
Counts mentionsInspects fields and schemasExecutes a published buying standard against the page
Reports who appearedFlags generic omissionsChecks every buyer requirement as an assertion
Produces one summary numberCannot execute a buyer taskPreserves the evidence trace
Cannot say why a journey failedCannot show model behaviorIsolates the store-controlled failure — and refuses to invent a fix when the cause is external
Cannot verify a correctionCannot rerun a specific failureReruns the identical test after the fix, and keeps it as a regression check

A machine can't act on a fact your store can't prove.

The engine improves by finding where it was wrong.

A test engine that is never measured against itself is a rubric with a user interface. So this one is run against large samples of real storefronts, and then every row it passed is read individually against that store's full page text — not sampled, not spot-checked. The passes that turn out to be wrong are counted, named, and published as an error bound with the sample and the method beside it.

Each confirmed wrong pass becomes a named defect class and a pinned case in an adversarial corpus, which fails in both directions: fix the defect and its case fails until the record is updated, reintroduce it and the case fails again. Defects we have chosen not to close are numbered, published with their measured cost, and left visible rather than quietly carried.

The bound has moved several times, and every move so far has come from the audit getting better rather than the engine getting worse — including one occasion when a figure we had published about ourselves turned out to measure what that audit had thought to look for, rather than the error rate. That correction is on the record too, at the version where it was made.

AI buyers treat your store like an API. We test it like one.

The standard is public before the test runs.

A buying standard is the set of questions a competent buyer in a category actually needs settled — and, for each one, what counts as evidence, what does not, and which surface decides when two of them disagree. AisleLens Coffee Standard v1.3 carries 42 such entries at a fixed version and content hash, so a result cites the exact contract it ran under and that citation still resolves a year later. Every entry is readable at its own URL, before you buy anything and before a test is run.

Ten of the 42 are executable against a public product page today. The other 32 are written down with the reason each one is not: 16 should be executable and the engine cannot reach them yet, each naming its own gap; 11 are real buyer questions that public data cannot adjudicate at all; and 5 are questions the engine can run and public data can settle, for which this standard has not yet written the binding and put it through the adversarial pass — recorded as unbound rather than quietly dropped.

We publish what we cannot test, and why.

Ten of forty-two is the honest ratio. A standard that listed only its own strengths would be marketing, and the second number is the one a merchant needs in order to know what a passing result did not cover.

A category standard is fitness-measured on its own category before we publish an error rate for it. It has been run against real coffee products on real storefronts, and every single requirement it passed was then read individually against that store's full page text — not sampled. The passes that turned out to be wrong are counted, and the measured upper bound on the error rate a coffee roaster should expect is published on the standard's own page with the sample size, the method and the defect classes behind it. We do not restate those figures here: this page cannot derive them, and a number typed beside a generated one is how a page goes quietly false.

The same discipline corrected a number we had published about ourselves. Our broad, non-category sample had been audited row by row and reported zero errors. Checking one defect class mechanically — a product identifier that is really the store's own internal id — found errors in that same sample that no reader could have caught, because that row shows the merchant no quote to be suspicious of. The figure had not been an error rate. It was a measurement of what that audit had thought to look for. The bound has moved three times since, each time because the audit got better, and every move is on the record.

Run it on one client.

The fastest way to judge this is to point it at a page you already know well and see whether the result matches what you would have said yourself. That takes one URL and no account. If you want to scope a category standard or a client engagement, say so and a person answers.

Run a real test

No account is needed for a test. A pilot is a conversation, not a checkout.

Questions

What is a buying standard?
The questions a competent buyer in a category actually needs settled, written down: each one with an assertion, the evidence that satisfies it, the evidence that specifically does not, and the rule that decides when two surfaces disagree. It is fixed at a version and a content hash, so the contract a result ran under can be cited and re-run exactly.
What is a "buyer task"?
A real shopping requirement, stated the way a customer would: "250 g of single-origin whole bean under £20, ground for espresso, dispatched this week." AisleLens turns each part into an assertion your store either proves or doesn't.
Is this SEO?
No. SEO is about which pages a search engine surfaces. This is about whether a machine acting for a buyer can settle specific requirements — a price cap, an ingredient claim, a variant in stock, a dispatch date — from what your store publishes. Different mechanism, different fix, testable outcome.
Is this an AI mention tracker?
No. Mention trackers count how often a brand appears and roll it into one number; that category is crowded and Shopify ships a free version. AisleLens publishes the standard for a category, executes it against your product pages, and reports each requirement as proven, not proven, or requires store access — with the evidence. External AI answers can seed our tests and appear in full diagnostics, as inputs rather than as the product.
Can you promise an AI assistant picks my product?
No, and anyone promising that is telling you something they cannot know. External AI systems update on their own schedule and weigh factors nobody controls. What we prove is narrower and real: a requirement a machine could not settle from your store is now settleable, and the identical test that failed now passes — reported honestly either way.
What if the problem isn't my store?
Then we tell you, and we don't sell you a fix. Some failures come from how external systems retrieve answers, or from third-party pages saying something wrong about you. The tool shows what it found and refuses to propose a store edit it can't justify.
Will you change my store without asking?
Never. Every change is proposed, previewed, approved by you, and reversible.

AI systems vary by model, prompt, time, and location. AisleLens reports what it tested and what it could verify from your store's own data. It makes no prediction about any external AI system, and is not affiliated with any AI provider.