AisleLens Coffee Standard v1.1

AisleLens Coffee Standard v1.1 — buyer questions, assertions and evidence rules for roasted coffee product pages

Standard
ALS-COFFEE
Version
1.1
Status
applied by author
Content hash
f8ec2780f60c38931913e5b6cd37506500c8462709209de7180ba6691d6137e7
Entries
42
Independently applied
No

Independently applied: no. We wrote this standard and we run it; no third party has applied it to a store. That is a limit on what a pass here is worth, and it is stated because a site about claim discipline cannot make its first unchecked claim about itself.

The full standard as JSON · Every source this standard is grounded in · A real result on a real store · How these are built and measured

Measured error

Every row this standard passed was audited individually against its full evidence. The bound is a 95% upper bound, cluster-adjusted at ICC 0.2 because pass rows are not independent — rows from one store share that store's copy conventions, and the bare rule of three would overstate the precision.

False-positive rate by sample
SampleStoresPass rows auditedConfirmed false positivesPoint estimate95% upper bound
Coffee category sample 7716210 6.17%12.78%
General DTC sample 16950918 3.54%7.80%

No comparison is drawn between these samples. General DTC sample is a FLOOR — one defect class was checked mechanically and the rest of that sample is unexamined — so it is not a like-for-like counterpart to a sample whose every passing row was read individually. A ratio between them would be arithmetic rather than a measurement. Every earlier version of the claim that a category sample and a general sample differ turned out to be measuring AUDIT DEPTH rather than category — 0.83% became 7.80% when a single defect class was re-checked mechanically — and at equal audit depth the two are statistically indistinguishable. The current measurement, with intervals, is on the current version's page.

Coffee category sample — method

Every pass_evidenced row audited individually against its FULL untruncated evidence, not the 180-character rendered quote. 162 of 162 adjudicated; the merge step refuses to produce a number if any row is unadjudicated. The 95% upper bound is a Poisson upper limit inflated by a design effect for clustering within store (ICC 0.2), because pass rows are not independent — rows from one store share that store's copy conventions.

What the coffee category sample's 10 wrong passes actually were
ClassnExampleAddressed by a guard
A brewing recipe or a caffeine dose per serving, read as the product's own net weight3 15g medium ground coffee 250ml water at 203°F no
A store-local Shopify product id or a byte-copy of the SKU, published in `mpn` and accepted as an external identifier3 "sku":"100754","mpn":"100754" — the store-local SKU the requirement explicitly excludes, because a SKU cannot match a product to an EXTERNAL catalogue no
The soil-science sense of `organic`, read as an organic certification claim2 organic soils and abundant rainfall no
`single-origin` inside a sentence describing a BLEND — the term is present and the sentence asserts the opposite of the requirement2 our Cold Brew Blend features a washed single-origin from Guatemala and a natural from Ethiopia no

Record: experiments/v3-2/CATEGORY_BOUND.md

General DTC sample — method

⚠️ THIS IS A FLOOR, NOT A COMPLETED AUDIT, AND IT CORRECTS A PREVIOUSLY PUBLISHED NUMBER. An earlier audit read all 507 rendered rows and confirmed ZERO false positives, which is where the 0.83% figure came from. A mechanical check of ONE defect class — `identifiers` rows passing on an `mpn` that is the store's own Shopify product id or a byte-copy of its SKU — found 18 in that same sample. The earlier audit could not have seen them: an identifier row renders NO QUOTE, so it looks identical whether the value is a real GS1 barcode or a number the store minted about itself. Only that one class has been re-checked, so the true figure is at least this and the two samples are NOT audited to the same depth.

Record: experiments/v3-2/CATEGORY_BOUND.md

What a passing claim row cannot rule out

What a passing claim row cannot rule out: on copy where the claim term sits next to a supplier, a farm, a region or a bundled item, the row establishes that the page states the term — not that the term was asserted of this product. Read the quoted sentence, which every passing row shows, and check what it is about.

We attack our own claim matcher with sentences written to break it. That measures CAPABILITY — what an adversary could do — and is deliberately not a measurement of what merchants write. Beside each capability figure is how often the same shape occurs in the sentences the engine actually rendered as evidence on real stores.

Attack shape, by capability and by real-copy frequency
What the sentence does Succeeds on sentences written to defeat it Occurs in real evidence sentences Status
The term is present, but the sentence does not assert it of this product — an invitation, a capability offer, a placeholder. 93.3% (252 of 270) 4.2% (3 of 71) not guarded — it does not occur
The property is described as past, future, conditional or merely possible rather than as holding now. 70.5% (425 of 603) 0% (0 of 71) not guarded — it does not occur
The term attaches to something other than the product — a supplier, a farm, a region, a bundled item, a practice. 48% (425 of 886) 15.5% (11 of 71) known limitation

The known limitation, and what it cost to try to close it

A guard for this axis was designed, implemented and measured in full. It closed 6 of its 8 real-copy targets and lost no true row on a 349-store replay. An independent adversarial pass — four attackers who wrote neither the guard nor its acceptance suite, 805 probes, every claimed regression re-executed by a refuter — confirmed 119 true statements it would have stopped reporting, against 6 defects closed. That is 19.8 true rows lost per defect closed, against a bar of 2.33–5.13. It was reverted and the limitation recorded as G-15.

The bar it had to clear was derived, not chosen: 14 honest carriers / 6 defects only this axis closes = 2.33; 41 / 8 = 5.13.

This table was completed on 2026-07-28 and deliberately NOT published until the open axis had either a fix or a recorded limitation — security-disclosure practice, decided before the outcome was known. Publishing selectively was considered and rejected: a table showing two axes at 'attacks well, occurs never' while omitting the third would read as a clean bill of health. It ships whole or not at all.

Measured 2026-07-28 against engine v2.4.0.

Measured 2026-07-27 against engine v2.0.0.

Worked example: the row that shows you nothing

“A product identifier is published in structured data” is the one requirement in this standard whose result renders NO QUOTE. It reports that your structured data publishes an identifier; it cannot show you the identifier, so the row reads the same whether the value is a real barcode or a number the store made up about itself.

The row's promise is that a machine buyer can match this exact product to an EXTERNAL catalogue. A store-local id cannot do that — it resolves inside one shop and nowhere else — which is why the requirement names the SKU as explicitly insufficient.

3 real stores, from the captured bytes of their own pages. All 3 passed this row when the example was built; 2 still do. One of them always deserved to.

sputnikcoffeecompany.com — 600160850004

Honest pass: A 12-digit UPC-A whose GS1 mod-10 check digit validates (recomputed: 4, published: 4), published as gtin12 and mirrored into mpn and the variant barcode. This resolves outside the store.

The engine passes this row today. Your structured data publishes an MPN (600160850004).

…ts/8oz-sputnik-coffee-ground-vacuum-tin?variant=46026182066429" } ], "gtin12": "600160850004", "productId": "600160850004", "brand": { "name": "Sputnik Coffee Company"…

From the page captured on 2026-07-27, at byte 26656. Published in: gtin12, mpn, barcode.

glowrecipe.com — 8079462006899

Was a false pass; now refused: The same value appears in 11 different fields on this page, among them Shopify's own `productId` and `data-product-id`, and alongside `gid://shopify/Product/8079462006899`. It is the store's internal product id, written into `mpn` by the theme. The store's real SKU is a different string entirely.

The engine no longer passes this row. The only identifier we can use on your product's own structured-data node is an MPN (8079462006899), and that is the id your own storefront uses for this product — it resolves to nothing outside your store. We read that node only, so a GTIN published beneath it on an offer or a variant is not counted here.

…ca548 data-product-gallery-projects="[]" data-collection-gallery-projects="[]" data-product-id=8079462006899 data-template-name="product" data-ot-ignore data-cookieconsent="ignore" > </sc…

From the page captured on 2026-07-26, at byte 261002. Published in: rid, mpn, id, data-product-id, ProductID, resourceId, productId, data-prodID, data-productid, product_id, content_ids.

www.stumptowncoffee.com — 100754

False pass, still live: The same six digits are published as BOTH `sku` and `mpn`, adjacent in one JSON-LD object. A SKU is the store's own part number — the requirement's own insufficient-evidence list names it — so copying it into `mpn` cannot make it resolve anywhere else.

The engine passes this row today. Your structured data publishes an MPN (100754).

…d": "https:\/\/www.stumptowncoffee.com\/products\/net-wrecker", "@type": "Product","sku": "100754","mpn": "100754","brand": { "@type": "Brand", "name": "Stumptown Coffee" }, "description":…

From the page captured on 2026-07-27, at byte 1030. Published in: sku, mpn.

This is the class an audit cannot see. There is no rendered evidence to be suspicious of, so a reader checking every row learns nothing from any of them. The store-local values above sat inside the general sample below — 509 pass rows, every one read individually — and a reader checking each rendered quote would have found none of them. It took one mechanical check against the captured bytes. The bound this page publishes has moved three times as the audit method improved, and each move was a measurement of what the previous audit had thought to look for.

1 of these 2 is closed and 1 is still live, and the difference is the honest part. The engine now refuses an MPN that is byte-identical to the storefront's own product object id. It does NOT refuse one that is a copy of the store's own SKU: the rule that would (“an MPN equal to the SKU”) was scored over every MPN-publishing product in the corpus at 0 true positives and 7 false, and the seven are the compliant case — a brand that manufactures what it sells legitimately uses one string for both. Closing three sentences is not closing a class.

Record: experiments/v3-2/CATEGORY_BOUND.md; extracted mechanically by experiments/v3-3/extract_identifier_example.mjs · engine behaviour re-executed 2026-07-28 at 132d085

Measured against predicted, entry by entry

The bands were written before this standard had ever run. They are kept exactly as authored and shown beside what actually happened, because a document that shows its own hypothesis failing is making a stronger case that its numbers are measured than one whose predictions all came true.

Fail rate per entry, over the products each entry was asked
EntryMeasuredPredicted VerdictAskedDiscriminates
FORMAT-001 73.7% 30-60% above band 99 yes
FORMAT-002 92.9% 35-70% above band 99 no
GRIND-001 84.8% 55-85% held 99 yes
GRIND-002 92.9% 60-85% above band 99 no
WEIGHT-001 49.0% 15-40% above band 100 yes
CERT-001 92.0% 30-70% above band 100 no
CERT-002 96.0% 40-75% above band 100 no
SOURCE-001 89.0% 40-75% above band 100 no
IDENT-001 94.7% 55-85% above band 76 no
DELIV-001 45.0% 50-80% below band 100 yes

1 of 10 predictions held. 8 of 10 came in above their band, which means those entries discriminate less than predicted, not more.

Discrimination and conformance are different questions, and only one of them is about whether an entry belongs here. 4 of 10 entries have a measured fail rate inside the 15-85% band where an answer carries information: FORMAT-001 at 73.7%, GRIND-001 at 84.8%, WEIGHT-001 at 49.0%, DELIV-001 at 45.0%. The others are failed by almost every store, so they separate nobody from anybody — and they are still legitimate requirements, because whether the page states the thing is still what a buyer needs settled. Nothing is deleted for discriminating poorly.

Executable (10)

This question is asked as a test and produces a result for a real page.

Back to contents ↑

Not discriminating (5)

This question is executable, but nearly every page answers it the same way — so the answer carries almost no information, and it is published rather than run.

Back to contents ↑

Advisory (11)

This question is worth asking but is not reducible to a check a page can settle. It is published as guidance, not as a test.

Back to contents ↑

Blocked (16)

This question matters to a buyer and CANNOT be answered from a public product page today. What would be required is stated in full.

Back to contents ↑