AisleLens Coffee Standard v1.2

AisleLens Coffee Standard v1.2 — buyer questions, assertions and evidence rules for roasted coffee product pages

Standard
ALS-COFFEE
Version
1.2
Status
applied by author
Content hash
fe199a864d3d4d565986851f9bfae9e108d55e4c86af18b1f8027f3d23486b58
Entries
42
Independently applied
No

Independently applied: no. We wrote this standard and we run it; no third party has applied it to a store. That is a limit on what a pass here is worth, and it is stated because a site about claim discipline cannot make its first unchecked claim about itself.

The full standard as JSON · Every source this standard is grounded in · A real result on a real store · How these are built and measured

Measured error

Every row this standard passed was audited individually against its full evidence. The bound is a 95% upper bound, cluster-adjusted at ICC 0.2 because pass rows are not independent — rows from one store share that store's copy conventions, and the bare rule of three would overstate the precision.

False-positive rate by sample
SampleStoresPass rows auditedConfirmed false positivesPoint estimate95% upper bound
ALS-COFFEE category sample 7716210 6.17%12.78%

ALS-COFFEE category sample — method

Every one of the 162 passing rows was adjudicated individually, against the FULL untruncated evidence rather than the rendered quote, 7 judgement calls were recorded and counted as PASSES, so they can be disputed rather than merely trusted.

Completion state: DEFECTS_FOUND. 10 confirmed false passes. A measurement that did not finish resolves to INCOMPLETE and may never be summed into a defect total or read as a pass — zero is the most dangerous number a broken instrument returns, because it is also what a healthy one returns.

Per-store rate: 12.99%, clustered at ICC 0.2 — pass rows are not independent, because rows from one store share that store's copy conventions.

The sample

100 products evaluated across 77 storefronts; selected by first applicable product; deduplicated by brand before capture; the applicability gate enforced; 2 excluded as out of category; 1 refused as unclassifiable; captured 2026-07-27.

The 162 audited rows are the pass rows of the same 100-product run, and they come from 77 of its 100 storefronts — the other 23 produced no passing row at all, which is why this store count is lower than the discrimination sample's and must never be substituted for it.

Record: experiments/v3-2/verdicts.json

Every one of the 10 wrong passes, individually

10 confirmed, against 10 recorded on the sample — the list accounts for all of them. Per entry: WEIGHT-001 3 · IDENT-001 3 · CERT-001 2 · SOURCE-001 2. 7 of 10 fire on vocabulary this category's pages contain and a general DTC page does not; 10 are pinned known gaps that no mechanism in the engine currently addresses.

Read all 10, with the store, the evidence the engine matched, and why it was wrong
Confirmed false passes
EntryStore and matched evidenceWhy it is wrongScopeStatus
WEIGHT-001 deathwishcoffee.com · A weight or measurement is stated in readable text · the engine matched: "a standard 6oz serving" The only numbers are a caffeine dose (210mg) and a brewed serving size ("a standard 6oz serving") — both explicitly not the product's own weight or capacity. category-specific no guard addresses it
WEIGHT-001 groundsforchange.com · A weight or measurement is stated in readable text · the engine matched: water/ice ounces, 1/4 cup of grounds, 200-205 degrees Every quantity in the evidence belongs to a Japanese-iced-coffee brewing recipe (water/ice ounces, 1/4 cup of grounds, 200-205 degrees), not to the product's own size. category-specific no guard addresses it
CERT-001 hydrangea.coffee · An organic claim is stated in readable text · the engine matched: soils are described as rich in organic matter "soils are described as rich in organic matter" is the soil-science sense of the word and makes no claim that the coffee is organic. category-specific no guard addresses it
IDENT-001 sightglasscoffee.com · A product identifier is published in structured data · the engine matched: it equals source_product_id/product.id on the page The JSON-LD mpn is "8631346921664", which is this store's own Shopify internal product ID (it equals source_product_id/product.id on the page), fails the GTIN-13 check digit, is identical across all eight variants, and no gtin field exists anywhere. general no guard addresses it
IDENT-001 www.lacolombe.com · A product identifier is published in structured data · the engine matched: it equals source_product_id/resourceId/product.id on the page The JSON-LD mpn is "7649496858737", the store's own Shopify internal product ID (it equals source_product_id/resourceId/product.id on the page), fails the GTIN-13 check digit, and no gtin field is published anywhere on the page. general no guard addresses it
IDENT-001 www.stumptowncoffee.com · A product identifier is published in structured data · the engine matched: 100754 The JSON-LD emits "sku":"100754","mpn":"100754" adjacent in the same object -- the mpn is a byte-identical echo of the store-local SKU, which the requirement explicitly excludes, and no gtin field exists. general no guard addresses it
SOURCE-001 blossomcoffeeroasters.com · A single-origin claim is stated in readable text · the engine matched: our Cold Brew Blend features a washed single-origin from Guatemala and a natural from Ethiopia The evidence says the opposite of the claim: 'our Cold Brew Blend features a washed single-origin from Guatemala and a natural from Ethiopia' describes a BLEND of two coffees — 'single-origin' modifies one component lot, not this product, which is by construction not single-origin. category-specific no guard addresses it
CERT-001 brashcoffee.com · An organic claim is stated in readable text · the engine matched: where organic soils and abundant rainfall have created an ideal terroir 'where organic soils and abundant rainfall have created an ideal terroir' uses 'organic' in the soil-science sense (organic matter in the soil at the Aquiares farm) — it makes no organic-certification or organic-coffee claim about this product. category-specific no guard addresses it
WEIGHT-001 myalmacoffee.com · A weight or measurement is stated in readable text · the engine matched: try this recipe : 15g medium ground coffee 250ml water at 203F 1:16.5 ratio The matched text is a BREWING RECIPE — "try this recipe : 15g medium ground coffee 250ml water at 203F 1:16.5 ratio" — a dose of grounds and a volume of water, not the product's own size; this is the exact excluded class, and the merchant is shown the recipe as proof. category-specific no guard addresses it
SOURCE-001 thewestbean.com · A single-origin claim is stated in readable text · the engine matched: Composed of three single-origin, estate grown beans "Composed of three single-origin, estate grown beans" says this product is a BLEND of three separate origins — the term is present but the sentence asserts the opposite of the requirement; the store does not state this coffee is single-origin. category-specific no guard addresses it

What this bound does NOT cover

  • This bounds SAMPLING error and nothing else. It does not cover a bias in the engine doing the measuring, and this run has one: the bounded semantic tier was disabled. Both of its paths are therefore absent — the missing GRANT path pushes a claim row's fail rate up and the missing VETO path pushes it down — and nothing here measures which dominates, so the tier is declared as two biases per affected entry rather than netted into one.
  • One product per storefront. 100 products over 100 distinct hosts means within-store copy variation is entirely unmeasured: nothing here says a second product from the same roaster would fare the same.
  • The comparison with the general DTC sample is a FLOOR, not a bound, and the two are NOT audited to the same depth. The general figure — 7.8% cluster-adjusted over 509 pass rows from 169 stores — rests on ONE defect class re-checked mechanically, after an audit that read every rendered row and confirmed zero. So the ratio between 7.8% and 12.78% is not itself a measurement; the DIRECTION is what the category rule rests on.
  • 10 confirmed false positives is a count of rows an auditor could recognise as wrong from the page text. A defect class that renders no quote is invisible to any audit that reads rendered evidence, however many rows it reads — which is how a previously published 0.83% general bound turned out to be 7.80%.
  • 7 rows were judgement calls counted as PASSES. Counting them as defects instead would raise every figure here; they are named in the audit record so the call can be disputed rather than merely trusted.
  • Fitness was measured on the ten entries that are executable. It says nothing about the 32 entries that are not run, including the 5 now marked `unbound`.

What a passing claim row cannot rule out

What a passing claim row cannot rule out: on copy where the claim term sits next to a supplier, a farm, a region or a bundled item, the row establishes that the page states the term — not that the term was asserted of this product. Read the quoted sentence, which every passing row shows, and check what it is about.

We attack our own claim matcher with sentences written to break it. That measures CAPABILITY — what an adversary could do — and is deliberately not a measurement of what merchants write. Beside each capability figure is how often the same shape occurs in the sentences the engine actually rendered as evidence on real stores.

Attack shape, by capability and by real-copy frequency
What the sentence does Succeeds on sentences written to defeat it Occurs in real evidence sentences Status
The term is present, but the sentence does not assert it of this product — an invitation, a capability offer, a placeholder. 93.3% (252 of 270) 4.2% (3 of 71) not guarded — it does not occur
The property is described as past, future, conditional or merely possible rather than as holding now. 70.5% (425 of 603) 0% (0 of 71) not guarded — it does not occur
The term attaches to something other than the product — a supplier, a farm, a region, a bundled item, a practice. 48% (425 of 886) 15.5% (11 of 71) known limitation

The known limitation, and what it cost to try to close it

A guard for this axis was designed, implemented and measured in full. It closed 6 of its 8 real-copy targets and lost no true row on a 349-store replay. An independent adversarial pass — four attackers who wrote neither the guard nor its acceptance suite, 805 probes, every claimed regression re-executed by a refuter — confirmed 119 true statements it would have stopped reporting, against 6 defects closed. That is 19.8 true rows lost per defect closed, against a bar of 2.33–5.13. It was reverted and the limitation recorded as G-15.

The bar it had to clear was derived, not chosen: 14 honest carriers / 6 defects only this axis closes = 2.33; 41 / 8 = 5.13.

This table was completed on 2026-07-28 and deliberately NOT published until the open axis had either a fix or a recorded limitation — security-disclosure practice, decided before the outcome was known. Publishing selectively was considered and rejected: a table showing two axes at 'attacks well, occurs never' while omitting the third would read as a clean bill of health. It ships whole or not at all.

Measured 2026-07-28 against engine v2.4.0.

Measured 2026-07-27.

Worked example: the row that shows you nothing

“A product identifier is published in structured data” is the one requirement in this standard whose result renders NO QUOTE. It reports that your structured data publishes an identifier; it cannot show you the identifier, so the row reads the same whether the value is a real barcode or a number the store made up about itself.

The row's promise is that a machine buyer can match this exact product to an EXTERNAL catalogue. A store-local id cannot do that — it resolves inside one shop and nowhere else — which is why the requirement names the SKU as explicitly insufficient.

3 real stores, from the captured bytes of their own pages. All 3 passed this row when the example was built; 2 still do. One of them always deserved to.

sputnikcoffeecompany.com — 600160850004

Honest pass: A 12-digit UPC-A whose GS1 mod-10 check digit validates (recomputed: 4, published: 4), published as gtin12 and mirrored into mpn and the variant barcode. This resolves outside the store.

The engine passes this row today. Your structured data publishes an MPN (600160850004).

…ts/8oz-sputnik-coffee-ground-vacuum-tin?variant=46026182066429" } ], "gtin12": "600160850004", "productId": "600160850004", "brand": { "name": "Sputnik Coffee Company"…

From the page captured on 2026-07-27, at byte 26656. Published in: gtin12, mpn, barcode.

glowrecipe.com — 8079462006899

Was a false pass; now refused: The same value appears in 11 different fields on this page, among them Shopify's own `productId` and `data-product-id`, and alongside `gid://shopify/Product/8079462006899`. It is the store's internal product id, written into `mpn` by the theme. The store's real SKU is a different string entirely.

The engine no longer passes this row. The only identifier we can use on your product's own structured-data node is an MPN (8079462006899), and that is the id your own storefront uses for this product — it resolves to nothing outside your store. We read that node only, so a GTIN published beneath it on an offer or a variant is not counted here.

…ca548 data-product-gallery-projects="[]" data-collection-gallery-projects="[]" data-product-id=8079462006899 data-template-name="product" data-ot-ignore data-cookieconsent="ignore" > </sc…

From the page captured on 2026-07-26, at byte 261002. Published in: rid, mpn, id, data-product-id, ProductID, resourceId, productId, data-prodID, data-productid, product_id, content_ids.

www.stumptowncoffee.com — 100754

False pass, still live: The same six digits are published as BOTH `sku` and `mpn`, adjacent in one JSON-LD object. A SKU is the store's own part number — the requirement's own insufficient-evidence list names it — so copying it into `mpn` cannot make it resolve anywhere else.

The engine passes this row today. Your structured data publishes an MPN (100754).

…d": "https:\/\/www.stumptowncoffee.com\/products\/net-wrecker", "@type": "Product","sku": "100754","mpn": "100754","brand": { "@type": "Brand", "name": "Stumptown Coffee" }, "description":…

From the page captured on 2026-07-27, at byte 1030. Published in: sku, mpn.

This is the class an audit cannot see. There is no rendered evidence to be suspicious of, so a reader checking every row learns nothing from any of them. The store-local values above sat inside the general sample below — a sample read row by row — and a reader checking each rendered quote would have found none of them. It took one mechanical check against the captured bytes. The bound this page publishes has moved three times as the audit method improved, and each move was a measurement of what the previous audit had thought to look for.

1 of these 2 is closed and 1 is still live, and the difference is the honest part. The engine now refuses an MPN that is byte-identical to the storefront's own product object id. It does NOT refuse one that is a copy of the store's own SKU: the rule that would (“an MPN equal to the SKU”) was scored over every MPN-publishing product in the corpus at 0 true positives and 7 false, and the seven are the compliant case — a brand that manufactures what it sells legitimately uses one string for both. Closing three sentences is not closing a class.

Record: experiments/v3-2/CATEGORY_BOUND.md; extracted mechanically by experiments/v3-3/extract_identifier_example.mjs · engine behaviour re-executed 2026-07-28 at 132d085

Measured discrimination, entry by entry

Each executable entry was run against the same recorded sample and its failure rate given a 95% interval, then compared with the 15-85% target band. THE INTERVAL DECIDES, not the point estimate — a rate outside the band whose interval straddles the edge has established nothing, and saying so is the difference between a measurement and a number.

Fail rate, 95% interval and verdict per entry
EntryFail rate95% interval Failed / adjudicatedVerdict
FORMAT-001 73.7% 64.3 – 81.4% 73 / 99 Discriminating
FORMAT-002 92.9% 86.1 – 96.5% 92 / 99 Not discriminating
GRIND-001 84.8% 76.5 – 90.6% 84 / 99 Discriminating
GRIND-002 92.9% 86.1 – 96.5% 92 / 99 Not discriminating
WEIGHT-001 49.0% 39.4 – 58.7% 49 / 100 Discriminating
CERT-001 92.0% 85.0 – 95.9% 92 / 100 Not discriminating
CERT-002 96.0% 90.2 – 98.4% 96 / 100 Not discriminating
SOURCE-001 89.0% 81.4 – 93.7% 89 / 100 Undecided
IDENT-001 94.7% 87.2 – 97.9% 72 / 76 Not discriminating
DELIV-001 60.8% 49.4 – 71.1% 45 / 74 (26 of 100 undecided by the engine, excluded) Discriminating

4 discriminating · 1 undecided · 5 not discriminating, of 10 measured entries.

Discriminating — the measured rate lies inside the target band, so the answer separates stores from one another.

Undecided — the rate is outside the band but its 95% interval is not — this measurement RAN AND DECIDED NOTHING, which is neither of the other two answers and must never be read as either. The entry stays executable and keeps accruing n.

Not discriminating — the whole 95% interval lies outside the target band, so almost every store answers the same way. This is evidence for a retirement decision and is not the decision.

1 of 10 decided nothing, and that is published rather than rounded to a decision. SOURCE-001 at 89.0% (95% interval 81.4–93.7%). An undecided result is "no difference detectable at this n", which is not "no difference".

5 of the 5 not-discriminating verdicts MAY NOT BE ACTED ON, and this document says so instead of working around it: FORMAT-002, GRIND-002, CERT-001, CERT-002, IDENT-001. A confidence interval bounds sampling error and nothing else, and each of these was measured with a known bias in the instrument pointing at the very band edge its interval cleared. None of the entries below has been retired. A measured verdict is EVIDENCE for a retirement decision; it is not the decision.

Executable (10)

This question is asked as a test and produces a result for a real page.

Back to contents ↑

Not yet bound (5)

The engine can run this kind of check and a public product page can settle it — but THIS standard has not yet written the binding or put the entry through its adversarial pass. The obstacle is unwritten work in this document, not the evidence and not the engine. Each one names the engine capability that fits.

  • ALS-COFFEE-1.2-PRICE-001 — What does this cost? No engine ReqKind fits as written. `price_under` exists but adjudicates a price against a cap the buyer supplies, and this entry asks only whether a price is published at all; the engine exports no `price_is_stated` kind. No binding and no adversarial pass have been authored for this entry in this standard, so it is not `executable`; the obstacle is unwritten work in this document, not the evidence and not the engine. Because no kind fits, this one needs an engine-gap proposal rather than a binding invented here.
  • ALS-COFFEE-1.2-STOCK-001 — Can I actually buy this right now? `in_stock` is a live engine ReqKind and public product data adjudicates it, so this is neither `advisory` (the evidence exists) nor `blocked` (the engine exists). No binding and no adversarial pass have been authored for this entry in this standard, so it is not `executable`; the obstacle is unwritten work in this document, not the evidence and not the engine.
  • ALS-COFFEE-1.2-TERMS-001 — Can I buy a single bag without signing up for a recurring subscription? `no_subscription` is a live engine ReqKind and public product data adjudicates it. It is absence-based, so a conformant result is `pass_no_blocking` and never `pass_evidenced` — a binding here would have to say so explicitly. No binding and no adversarial pass have been authored for this entry in this standard, so it is not `executable`; the obstacle is unwritten work in this document, not the evidence and not the engine.
  • ALS-COFFEE-1.2-DECAF-004 — Can I buy a decaf version of this same coffee? No engine ReqKind fits: this entry adjudicates a statement about a DIFFERENT product from the one under test — whether a decaf counterpart exists elsewhere in the catalogue — and every kind the engine exports reads the product in front of it. No binding and no adversarial pass have been authored for this entry in this standard, so it is not `executable`; the obstacle is unwritten work in this document, not the evidence and not the engine. Because no kind fits, this one needs an engine-gap proposal rather than a binding invented here.
  • ALS-COFFEE-1.2-DIET-001 — Is this coffee vegan and gluten-free? `claim` is a live engine ReqKind and `vegan` and `gluten_free` are both live claim keys in the engine's own claim table, so public product copy adjudicates it. No binding and no adversarial pass have been authored for this entry in this standard, so it is not `executable`; the obstacle is unwritten work in this document, not the evidence and not the engine.

Back to contents ↑

Advisory (11)

This question is worth asking but is not reducible to a check a page can settle. It is published as guidance, not as a test.

Back to contents ↑

Blocked (16)

This question matters to a buyer and CANNOT be answered from a public product page today. What would be required is stated in full.

Back to contents ↑