PUBLISHED BUYING STANDARDS · EXECUTABLE TESTS

The questions a competent buyer asks, written as executable tests.

AisleLens publishes versioned buying standards — one per category, fixed at a version and a content hash so a result can cite the exact contract that produced it — and runs them against your real product pages. AI buyers treat your store like an API. We test it like one.

One product. One buyer task. Proven, not proven, or requires store access — with the evidence.

Read Coffee Standard v1.3 → · All published standards · Methodology · See an example test →

The standard is public before the test runs.

A buying standard is the set of questions a competent buyer in a category actually needs settled — and, for each one, what counts as evidence, what does not, and which surface decides when two of them disagree. AisleLens Coffee Standard v1.3 carries 42 such entries at a fixed version and content hash, so a result cites the exact contract it ran under and that citation still resolves a year later. Every entry is readable at its own URL, before you buy anything and before a test is run.

Ten of the 42 are executable against a public product page today. The other 32 are written down with the reason each one is not: 16 should be executable and the engine cannot reach them yet, each naming its own gap; 11 are real buyer questions that public data cannot adjudicate at all; and 5 are questions the engine can run and public data can settle, for which this standard has not yet written the binding and put it through the adversarial pass — recorded as unbound rather than quietly dropped.

We publish what we cannot test, and why.

Ten of forty-two is the honest ratio. A standard that listed only its own strengths would be marketing, and the second number is the one a merchant needs in order to know what a passing result did not cover.

A category standard is fitness-measured on its own category before we publish an error rate for it. It has been run against 100 real coffee products across 77 storefronts, and every single requirement it passed was then read individually against that store's full page text — not sampled. The passes that turned out to be wrong are counted, and the measured upper bound on the error rate a coffee roaster should expect is published on the standard's own page with the method and the defect classes behind it. We do not restate those figures here: this page cannot derive them, and a number typed beside a generated one is how a page goes quietly false.

The same discipline corrected a number we had published about ourselves. Our broad, non-category sample had been audited row by row and reported zero errors. Checking one defect class mechanically — a product identifier that is really the store's own internal id — found errors in that same sample that no reader could have caught, because that row shows the merchant no quote to be suspicious of. The figure had not been an error rate. It was a measurement of what that audit had thought to look for. The bound has moved three times since, each time because the audit got better, and every move is on the record.

A summary number tells you that something moved. A test tells you what broke.

Mention monitoringReadiness checklistsAisleLens
Counts mentionsInspects fields and schemasExecutes a published buying standard against the page
Reports who appearedFlags generic omissionsChecks every buyer requirement as an assertion
Produces one summary numberCannot execute a buyer taskPreserves the evidence trace
Cannot say why a journey failedCannot show model behaviorIsolates the store-controlled failure — and refuses to invent a fix when the cause is external
Cannot verify a correctionCannot rerun a specific failureReruns the identical test after the fix, and keeps it as a regression check

A machine can't act on a fact your store can't prove.

How testing works

  1. Take the standard, or state the task A published category standard supplies the requirements and the evidence rules. Outside a published category, state the buying task yourself: attributes, price, variant, inventory, subscription terms, delivery, returns, compatibility.
  2. Execute the test Each requirement becomes a testable assertion, run against the evidence your store actually exposes — using explicit retrieval tools and traceable evidence checks.
  3. Trace the failure Which requirement failed, which surfaces were checked, what evidence was found, and where the test stopped rather than guessed.
  4. Correct what your store controls Approve a targeted, reversible change linked to the failed assertion. When the cause is external, AisleLens says so and does not manufacture a store fix.
  5. Rerun and retain The identical test runs again — pass or fail, reported either way. Passing tests become permanent regression checks against future catalog, policy, variant, and model changes.

Real-world signals can become tests. AisleLens converts shopping questions observed across external AI systems into executable store tests — including which requirements those systems emphasized and which stores they surfaced. Observed behavior seeds the tests; the tests do the proving.

One failed test. One isolated cause. One verified rerun.

  • Before the fix: 0 of 4 test runs passed — the required claim could not be verified from any store surface.
  • After one approved, reversible correction: 4 of 4 passed. Same test, same models, versions pinned.
  • Unsupported evidence credited: zero — every claim in every run traces to retrieved evidence.

This is a controlled technical validation on a Shopify development store, labeled as such. It is not a merchant case, and nothing on this page presents it as one.

Questions

What is a buying standard?
The questions a competent buyer in a category actually needs settled, written down: each one with an assertion, the evidence that satisfies it, the evidence that specifically does not, and the rule that decides when two surfaces disagree. It is fixed at a version and a content hash, so the contract a result ran under can be cited and re-run exactly.
What is a "buyer task"?
A real shopping requirement, stated the way a customer would: "250 g of single-origin whole bean under £20, ground for espresso, dispatched this week." AisleLens turns each part into an assertion your store either proves or doesn't.
Is this SEO?
No. SEO is about which pages a search engine surfaces. This is about whether a machine acting for a buyer can settle specific requirements — a price cap, an ingredient claim, a variant in stock, a dispatch date — from what your store publishes. Different mechanism, different fix, testable outcome.
Is this an AI mention tracker?
No. Mention trackers count how often a brand appears and roll it into one number; that category is crowded and Shopify ships a free version. AisleLens publishes the standard for a category, executes it against your product pages, and reports each requirement as proven, not proven, or requires store access — with the evidence. External AI answers can seed our tests and appear in full diagnostics, as inputs rather than as the product.
Can you promise an AI assistant picks my product?
No, and anyone promising that is telling you something they cannot know. External AI systems update on their own schedule and weigh factors nobody controls. What we prove is narrower and real: a requirement a machine could not settle from your store is now settleable, and the identical test that failed now passes — reported honestly either way.
What if the problem isn't my store?
Then we tell you, and we don't sell you a fix. Some failures come from how external systems retrieve answers, or from third-party pages saying something wrong about you. The tool shows what it found and refuses to propose a store edit it can't justify.
Will you change my store without asking?
Never. Every change is proposed, previewed, approved by you, and reversible.

AI systems vary by model, prompt, time, and location. AisleLens reports what it tested and what it could verify from your store's own data. It makes no prediction about any external AI system, and is not affiliated with any AI provider.