AisleLens Coffee Standard v1.3

AisleLens Coffee Standard v1.3 — buyer questions, assertions and evidence rules for roasted coffee product pages

Standard
ALS-COFFEE
Version
1.3
Status
applied by author
Content hash
ba2050578ed0274885fd6213967c230b2a57dd2b7c1d3fba8c5e1633027d4cf7
Entries
42
Independently applied
No

Independently applied: no. We wrote this standard and we run it; no third party has applied it to a store. That is a limit on what a pass here is worth, and it is stated because a site about claim discipline cannot make its first unchecked claim about itself.

The full standard as JSON · Every source this standard is grounded in · A real result on a real store · How these are built and measured

Measured error

Every row this standard passed was audited individually against its full evidence. The bound is a 95% upper bound, cluster-adjusted at ICC 0.2 because pass rows are not independent — rows from one store share that store's copy conventions, and the bare rule of three would overstate the precision.

This supersedes the measurement inside the document, and the document is not edited. category_fitness in standard.json records 10 confirmed false passes over 162 audited rows — a 12.78% bound, measured 2026-07-27. Its bytes are what a citation resolves through, so a measurement taken after publication goes beside the document rather than into it. v1.3's in-document block was carried across from v1.2 unchanged and records the engine as it was before rule D: 162 pass rows, 10 confirmed, 12.78% cluster-adjusted. Rule D disqualifies an MPN that is the storefront's own product object id, which removes two of those ten and the two rows that carried them. The document is not edited; this sidecar is the later measurement.

False-positive rate by sample
SampleStoresPass rows auditedConfirmed false positivesPoint estimate95% upper bound
Coffee category sample 771607 4.38%9.99%
General DTC sample 16948311 2.28%5.17%

No difference is stated between these samples, because their intervals overlap. Coffee category sample is 4.38% (95% 2.14–8.75%) and General DTC sample is 2.28% (95% 1.28–4.03%). Both have now had every passing row adjudicated individually, which is the first time the two have been audited to the same depth — and at equal depth they are statistically indistinguishable. Every earlier version of the claim that a category sample and a general sample differ was measuring AUDIT DEPTH rather than category: a general figure of 0.83% became 7.80% when one defect class was checked mechanically, and the general sample's remaining rate was unmeasured until every row was read. A ratio between overlapping intervals would be arithmetic, so none is stated.

Do the kinds differ from each other?

Two kinds are reported as different only when their 95% intervals do not overlap. That is the conservative test, and it is the one that matters: a spread stated on point estimates is a spread stated on nothing.

  • coffee: 1 of 10 pairs separate, the widest by 0.45 percentage points (variant_option vs claim).
  • general: 1 of 21 pairs separate, the widest by 0.17 percentage points (price_under vs in_stock).

So the decomposition does not support a spread at these sample sizes. It is published as counts — which kind the known errors are actually in — rather than as a set of rates, because a table of six percentages invites a comparison the intervals refuse.

Coffee category sample — method

Every one of the 162 pass rows the published measurement audited was individually adjudicated against its FULL untruncated evidence. Those adjudications are carried forward here unchanged; what moved is which rows the engine still passes. Rule D removed 2 of them, both adjudicated as confirmed false positives, so both the numerator AND the denominator fall — a correction that removed only the numerator would flatter the rate. The 95% upper bound is a Poisson upper limit inflated by a design effect for clustering within store (ICC 0.2), because rows from one store share that store's copy conventions.

Completion state: DEFECTS_FOUND. 7 confirmed false passes. A measurement that did not finish resolves to INCOMPLETE and may never be summed into a defect total or read as a pass — zero is the most dangerous number a broken instrument returns, because it is also what a healthy one returns.

Per-store rate: 9.09%, clustered at ICC 0.2 — pass rows are not independent, because rows from one store share that store's copy conventions.

The sample

100 products evaluated across 77 storefronts; selected by first applicable product; deduplicated by brand before capture; the applicability gate enforced; 2 excluded as out of category; 1 refused as unclassifiable; captured 2026-07-27.

The 103-brand run is deduplicated on registrable domain, and the audit's store count is lower than the run's because 23 storefronts produced no passing row at all. ⚠️ A DUPLICATION IN THE WIDER CORPUS THAT DOES NOT REACH THIS NUMBER, recorded so a future bound does not inherit it: the three snapshot sets this project holds (172 general, 44 coffee, 122 coffee) total 338 files but only 334 registrable domains and 335 distinct product URLs — `onyxcoffeelab.com` and `vervecoffee.com` are the SAME product in the general and coffee sets, and `deathwishcoffee.com` appears at both apex and www within one coffee set. Neither published sample pools across sets and both are internally deduplicated (general: 172 files, 172 domains, 172 URLs; this audit: 77 hosts, 77 domains, 77 URLs), so no figure here moves. Any bound computed over the UNION of the sets must dedupe first: two files of one product are perfectly correlated, not merely clustered, and would inflate n while adding no information.

Record: experiments/v3-5/publish/out/std_head_v10.jsonl

Coffee category sample — the bound decomposed by requirement kind
KindEntriesPass rowsStoresConfirmedBorderlinePoint95% intervalNaive P95 upperCluster-adjustedState
variant_option 45528 0 1 0.00% 0.0–6.5% 5.45% 6.50% VERIFIED_CLEAN
attribute 15151 3 0 5.88% 2.0–15.9% 15.20% refused DEFECTS_FOUND
delivery 12929 0 1 0.00% 0.0–11.7% 10.33% refused VERIFIED_CLEAN
claim 32320 4 5 17.39% 7.0–37.1% ⚠️ 39.80% 40.99% DEFECTS_FOUND
identifiers 122 0 0 0.00% 0.0–65.8% ⚠️ refused refused VERIFIED_CLEAN

Refused: the Poisson-upper/n form returns 149.79% at x=0, n=2 — above 100 it is not a rate. Read the Wilson interval instead; the published blended bound uses this form legitimately only because its n is large enough for the quotient to stay below 1.

Refused: rows_per_store == 1 — one product per store and one row per store in this kind, so there is no within-store correlation to adjust for. The naive figure IS the clustered figure.

Refused: the figure it would inflate is itself refused

Every one of the 7 wrong passes, individually

7 confirmed, against 7 recorded on the sample — the list accounts for all of them. Per entry: WEIGHT-001 3 · CERT-001 2 · SOURCE-001 2. 7 of 7 fire on vocabulary this category's pages contain and a general DTC page does not; 7 are pinned known gaps that no mechanism in the engine currently addresses.

Read all 7, with the store, the evidence the engine matched, and why it was wrong
Confirmed false passes
EntryStore and matched evidenceWhy it is wrongScopeStatus
WEIGHT-001 deathwishcoffee.com — CAFFEINE CONTENT: Power Surge Single-Serve Pods contain approximately 210mg of caffeine (20% more than our Dark Roast) based on a standard 6oz serving when brewed following… The only numbers are a caffeine dose (210mg) and a brewed serving size ("a standard 6oz serving") — both explicitly not the product's own weight or capacity. category-specific no guard addresses it
WEIGHT-001 groundsforchange.com — The essential concept is to brew double-strength coffee directly onto ice: What you need Drip cone with filter Equal amounts (by weight) filtered hot water (200-205 degrees) and… Every quantity in the evidence belongs to a Japanese-iced-coffee brewing recipe (water/ice ounces, 1/4 cup of grounds, 200-205 degrees), not to the product's own size. category-specific no guard addresses it
CERT-001 hydrangea.coffee — The farm’s volcanic soils are described as rich in organic matter, while narrow canyons channel warm winds through the surrounding landscape—the feature that gave Brisa Eterna its… "soils are described as rich in organic matter" is the soil-science sense of the word and makes no claim that the coffee is organic. category-specific no guard addresses it
SOURCE-001 blossomcoffeeroasters.com — Dark roasted to perfection, our Cold Brew Blend features a washed single-origin from Guatemala and a natural from Ethiopia, creating the perfect balance between clean and sweet. The evidence says the opposite of the claim: 'our Cold Brew Blend features a washed single-origin from Guatemala and a natural from Ethiopia' describes a BLEND of two coffees — 'single-origin' modifies one component lot, not this product, which is by construction not single-origin. category-specific no guard addresses it
CERT-001 brashcoffee.com — Since 1890, coffee has flourished at Aquiares, where organic soils and abundant rainfall have created an ideal terroir for aromatic coffee. 'where organic soils and abundant rainfall have created an ideal terroir' uses 'organic' in the soil-science sense (organic matter in the soil at the Aquiares farm) — it makes no organic-certification or organic-coffee claim about this product. category-specific no guard addresses it
WEIGHT-001 myalmacoffee.com — For a fun and refreshing cup, try this recipe : 15g medium ground coffee 250ml water at 203°F 1:16.5 ratio From the farmers who nurtured this coffee in Ethiopia to our roasters… The matched text is a BREWING RECIPE — "try this recipe : 15g medium ground coffee 250ml water at 203F 1:16.5 ratio" — a dose of grounds and a volume of water, not the product's own size; this is the exact excluded class, and the merchant is shown the recipe as proof. category-specific no guard addresses it
SOURCE-001 thewestbean.com — Composed of three single-origin, estate grown beans. "Composed of three single-origin, estate grown beans" says this product is a BLEND of three separate origins — the term is present but the sentence asserts the opposite of the requirement; the store does not state this coffee is single-origin. category-specific no guard addresses it

What this bound does NOT cover

  • ⚠️ THE STUMPTOWN ROW IS RE-SCORED AS A TRUE PASS, AND THIS SUPERSEDES THIS SIDECAR'S OWN EARLIER READING. www.stumptowncoffee.com publishes "sku":"100754","mpn":"100754". The earlier reading counted it as a surviving defect the engine deliberately does not catch. Scored against ALS-COFFEE v1.3's ACTUAL TEXT it is not a defect at all, so there is nothing for the engine to catch and counting it inflated the bound. The deciding sentence is IDENT-001's own adversarial.residual_risk clause (2), verbatim: "It disqualifies a value for being the seller's own object id, not for being seller-private in general: a stock code that is neither a placeholder nor the storefront's key is outside this clause." accepted_evidence and insufficient_evidence scope identically. Checked against the bytes rather than against that reasoning: the storefront's own product key is 9516469289128 and its variant key is 55754751967400 (engine shopifyStorefrontObjectId and meta.product.id agree); 100754 is the merchant's SKU. The third key carrying the value is the analytics defaultVariant.id / items[].id, whose conventional value in a GA-shaped payload IS the SKU — reading that as "the storefront's key" would disqualify every published SKU, which is the reading residual_risk (2) forecloses and the rule (mpn === sku) already measured at 0 true positives and 7 false.
  • ⚠️ A TENSION INSIDE THE ENTRY, STATED RATHER THAN RESOLVED SILENTLY. IDENT-001's insufficient_evidence clause for the MPN field reasons field-agnostically in its why_not — "A seller-private stock code is not a global identifier ... That reason is field-agnostic" — which read alone would disqualify a SKU echo. Its `form` and its residual_risk do not. `form` is the operative text an evaluator matches against and why_not is justification prose, so the rule governs; but the document owes a narrower why_not or a wider form, and until it has one this row is decided by a reading rather than by a match.
  • The bound covers the engine's false PASSES only. False fails are not in it, and rule D creates them: six real merchants in the captured corpus publish a check-digit-valid GTIN on a Product node the extractor does not select, and are now told they publish no usable identifier (ENGINE_GAPS P-12).
  • The comparison with the general DTC sample is now like for like for the first time — both samples have had every passing row adjudicated individually — and the two are STATISTICALLY INDISTINGUISHABLE. See cross_sample_comparison. No ratio is stated between them, because their intervals overlap.

General DTC sample — method

Every one of the 488 pass rows was adjudicated individually against its FULL untruncated evidence, in 13 kind-scoped batches, with every row assigned exactly once by (host,label) and the merge refusing to produce any number if a single row were unadjudicated. Each of the 18 confirmed defects was then RE-EXECUTED against the raw captured bytes by a mechanical rule that does not read the adjudicator's sentence, because auditor prose is a candidate and never a verdict: 18 candidates, 18 confirmed, 0 refuted, 0 uncheckable, with a two-sided canary of 108 executions of the same checks against adjudicated TRUE passes and 0 false alarms. That canary caught a defect in one of the checks on its first run — an availability check that read only the /products/{handle}.json tier fired on four true passes whose available flags live in the .js tier — so four extra defects would have been published without it. The 95% upper bound is a Poisson upper limit inflated by a design effect for clustering within store (ICC 0.2); a Wilson interval is published beside it because the per-kind cells are where small n arrives and the Poisson/n form stops being a probability there.

Completion state: DEFECTS_FOUND. 11 confirmed false passes. A measurement that did not finish resolves to INCOMPLETE and may never be summed into a defect total or read as a pass — zero is the most dangerous number a broken instrument returns, because it is also what a healthy one returns.

Per-store rate: 5.92%, clustered at ICC 0.2 — pass rows are not independent, because rows from one store share that store's copy conventions.

The sample

172 products evaluated across 169 storefronts; selected by one product per store; deduplicated by brand before capture; captured 2026-07-26.

172 files, 172 registrable domains, 172 distinct product URLs — checked before counting, because ENGINE_GAPS P-16 records that the wider 338-file corpus holds only 334 merchants. This sample is internally clean and pools with nothing. `stores` is 169 rather than 172 because three storefronts produced no passing row at all; 17 of those 169 carry at least one confirmed defect.

Record: experiments/v3-7/general_head.jsonl (replay of experiments/v2-9/snaps at 7085b34)

The per-kind decomposition for this sample is withheld. Its cells sum to 488 pass rows and 18 confirmed false positives, against this sample's 483 and 11 — they were measured before a later engine change and have not been re-derived. The counts could be corrected arithmetically; the per-cell intervals could not, and a table of repaired counts beside stale intervals would read as measured when it is not. The headline bound above is unaffected: it is computed from the audited rows directly, not from these cells.

What the general dtc sample's 11 wrong passes actually were
ClassnExampleAddressed by a guard
$0.00 treated as a price6 knifewear.com — (the row renders no quote) Lowest readable price is $0.00. no
A non-USD price rendered with a US dollar sign5 gardenerskit.com — (the row renders no quote) Lowest readable price is $75.00. no
Integer CENTS read as dollars — a factor of 1002 levainbakery.com — (the row renders no quote) Lowest readable price is $1000.00. no
A MISSING availability field defaulted to purchasable2 kytebaby.com — (the row renders no quote) A "Gilmore Girls" variant is listed and purchasable. no
A SIBLING product's size read as this product's1 askinosie.com — Try our 1/2lb. no
"Lowest readable price" is the page's MAXIMUM1 fieldcompany.com — (the row renders no quote) Lowest readable price is $135.00. no
Structured data says in stock; the store's own variant data says otherwise1 lesserevil.com — (the row renders no quote) Your structured data marks this product in stock. no

Every one of the 18 wrong passes, individually

18 confirmed, against 11 recorded on the sample — THESE DO NOT RECONCILE, and the shorter list is the incomplete one. Per entry: default-requirement/price_under 14 · default-requirement/in_stock 2 · default-requirement/attribute 1 · default-requirement/variant_option 1. 0 of 18 fire on vocabulary this category's pages contain and a general DTC page does not; 18 are pinned known gaps that no mechanism in the engine currently addresses.

Read all 18, with the store, the evidence the engine matched, and why it was wrong
Confirmed false passes
EntryStore and matched evidenceWhy it is wrongScopeStatus
default-requirement/attribute askinosie.com — Try our 1/2lb. The only quantity in the quoted structured-data evidence is the size of a DIFFERENT SKU: the raw JSON-LD description reads "Looking for a smaller size? Try our 1/2lb. Cocoa Nib Pouch.", an anchor to /products/mababu-tanzania-single-origin-roasted-cocoa-nibs-1-2-lb, while this product is the 1lb pouch. The merchant is shown "Try our 1/2lb." as proof of their own product's measurement. general no guard addresses it
default-requirement/price_under branchbasics.com — (the row renders no quote) Lowest readable price is $0.00. 'Lowest readable price is $0.00.' is a null treated as a number. The only price evidence is a Shopify variant price of '0.00' on a hidden fulfillment stub: title 'Concentrate 2PK - (Amazon Replacement)', tags 'os-3EiSwCJzni, unsearchable', empty body_html, available:false, and no JSON-LD Product/Offer on the page at all. Nothing establishes that this product is sold for under $10; the store publishes no price for it. general no guard addresses it
default-requirement/price_under fieldcompany.com — (the row renders no quote) Lowest readable price is $135.00. The sentence asserts a superlative the page's own bytes refute: the JSON-LD publishes a single Offer at 135.0 USD (no AggregateOffer, no lowPrice), but the Shopify analytics bootstrap on the same HTML lists 9 variants of 'Leather Oven Mitts, Factory Second' at 7900/7900/7900/9400x5/13500 cents — three Black Suede sizes at $79.00. $135.00 is the MAXIMUM variant price, rendered as 'Lowest readable price', and the derived cap should have been $80 not $140. general no guard addresses it
default-requirement/price_under gardenerskit.com — (the row renders no quote) Lowest readable price is $75.00. The number 75 is CANADIAN dollars, not US: the page's JSON-LD offer is "price": 75.0, "priceCurrency": "CAD", Shopify.currency = {"active":"CAD","rate":"1.0"}, Shopify.country = "CA", shop gardeners-kit.myshopify.com, paymentSettings CAD. The sentence renders C$75 as "$75.00" (≈US$55), so the figure the merchant reads is false. general no guard addresses it
default-requirement/price_under hismileteeth.com — (the row renders no quote) Lowest readable price is $34.99. This is the AUSTRALIAN storefront: shop hismile.myshopify.com, Shopify.currency = {"active":"AUD","rate":"1.0"}, Shopify.country = "AU", paymentSettings currencyCode AUD, and the analytics track call is {"currency":"AUD"}. The .json price 34.99 is A$34.99 (≈US$23) rendered to the merchant as "$34.99". general no guard addresses it
default-requirement/price_under knifewear.com — (the row renders no quote) Lowest readable price is $0.00. $0.00 is not a price. The .json shows product_type "Referral", body_html "", tags "meta-no-description, meta-no-image", and one variant priced "0.00" with SKU "ph-46838681075886" (a placeholder id) — a store-internal Calgary Farmers Market referral record. The row proves a price cap from the absence of a price, and additionally the shop's own currency is CAD (paymentSettings CAD, countryCode CA) served through an /en-us USD market at rate 0.709698. general no guard addresses it
default-requirement/price_under kosas.com — (the row renders no quote) Lowest readable price is $0.00. $0.00 is not a price: the .json shows product_type "GWP", template_suffix "gwp", and both shade variants priced "0.00" with compare_at_price "26.00" (Klaviyo renders Price "$0.00" / CompareAtPrice "$26.00"). This is a gift-with-purchase promo record granted on a qualifying order, so a price cap is being proven from a zero placeholder rather than from a price. general no guard addresses it
default-requirement/variant_option kytebaby.com — (the row renders no quote) A "Gilmore Girls" variant is listed and purchasable. The row says the variant is "purchasable" but every surface in the capture says the opposite: the page's own embedded variant records report all six "Gilmore Girls / …" variants as "available":false, and the page emits six JSON-LD offers (skus 1615GG1-1615GG6) all with availability schema.org/OutOfStock. The engine's product source here was /products/womens-short-sleeve-pajama-set-in-gilmore-girls.json, which contains no "available" key on any variant (verified: no "available" substring in the whole body), so the missing field was defaulted to true — 6/6 matched variants flagged available in the batch, 0/6 available per the page. Nothing on this product is buyable. general no guard addresses it
default-requirement/in_stock kytebaby.com — (the row renders no quote) At least one variant is listed as purchasable. The row asserts "At least one variant is listed as purchasable" but no public surface lists any variant as purchasable: the captured /products/womens-short-sleeve-pajama-set-in-gilmore-girls.json variant objects carry no `available` key at all (verified field-by-field on all 6), robots.txt blocks the .js endpoint with `Disallow: /*.js$` so the availability tier was never reached, and the engine's `available: v.available !== false` turned that absence into purchasable. The only availability the page actually states is 6 occurrences of "availability":"http://schema.org/OutOfStock" — one per size, with zero InStock anywhere — and 32 inline "available":false flags for this product with 0 true. This should have been requires_store_access. general no guard addresses it
default-requirement/in_stock lesserevil.com — (the row renders no quote) Your structured data marks this product in stock. The product is sold out and the row tells the merchant it is in stock. The page's inline Shopify product object states "available":false at both product level (price_min/price_max 929) and at its single default variant, 4 occurrences with 0 true. The page also publishes three mutually contradictory LD Product nodes for this product: two say InStock and a third — carrying the IDENTICAL @id "/products/cowboy-cheddar-cheezmos-snack-box#product" as the first — says "availability":"OutOfStock". The engine read the first node and rendered a pass; no .json or .js tier was fetched, so nothing arbitrated the conflict. An InStock marking that the same page contradicts twice does not establish "In stock and purchasable". general no guard addresses it
default-requirement/price_under levainbakery.com — (the row renders no quote) Lowest readable price is $1000.00. The mug costs $10.00, not $1000.00. /products/dad-mug.json says "price":"10.00", the web-pixel payload says {"amount":10.0,"currencyCode":"USD"} and Klaviyo says Price "$10.00"; the engine took the cents integer 1000 from /products/dad-mug.js (and the analytics bootstrap var meta …"price":1000) and priceToUsd's cents-guard only divides when p > 1000, so exactly 1000 cents falls through. Both the label "Price under $1005" and the sentence "$1000.00" are off by 100×. general no guard addresses it
default-requirement/price_under missoma.com — (the row renders no quote) Lowest readable price is $135.00. The number is POUNDS: the JSON-LD offer is "price":"135.00","priceCurrency":"GBP" (shipping also quoted in GBP), shop missoma-store.myshopify.com, Shopify.currency = {"active":"GBP","rate":"1.0"}, Shopify.country = "GB", paymentSettings GBP. £135 is roughly US$170, so rendering it as "$135.00" makes the row's own conclusion — under the $140 cap — false, not merely mislabelled. general no guard addresses it
default-requirement/price_under mustardmade.com — (the row renders no quote) Lowest readable price is $32.00. This is the AU storefront: JSON-LD offer "price" : 32.0, "priceCurrency" : "AUD", shop mustard-made.myshopify.com, theme literally named "[LIVE] Mustard Template - AU", Shopify.currency = {"active":"AUD","rate":"1.0"}, Shopify.country = "AU", paymentSettings AUD. A$32 (≈US$21) is rendered to the merchant as "$32.00". general no guard addresses it
default-requirement/price_under organicbasics.com — (the row renders no quote) Lowest readable price is $50.00. The captured storefront is the German/EUR market — Shopify.country = "DE", Shopify.currency = {"active":"EUR"}, og:price:amount "50,00" with og:price:currency EUR, and both JSON-LD Product nodes carry priceCurrency EUR — so the figure is €50, rendered by the engine as "$50.00" against a "$55" cap. general no guard addresses it
default-requirement/price_under partakefoods.com — (the row renders no quote) Lowest readable price is $0.00. $0.00 is a zeroed-out record, not a price: this is a real retail cookie SKU (60-07008) marked OutOfStock, and the same page's recommendation payload lists Partake's other cookies at real prices ($7.49, $16.99) alongside a cluster of retired SKUs all at 0.0 USD. The merchant is told a $0.00 price fact was proven for a product that is simply no longer priced. general no guard addresses it
default-requirement/price_under richer-poorer.com — (the row renders no quote) Lowest readable price is $300.00. The rendered sentence is false by a factor of 100. product.js gives price 300 / price_max 800 in CENTS and product.json confirms "3.00" and "8.00" (compare_at 12.00) — these socks cost $3.00. priceToUsd's cents-guard only divides when the integer exceeds 1000, so 300 and 800 passed through untouched, the label cap was generated from the bogus $300, and the merchant reads "Lowest readable price is $300.00" for a $3 pair of socks. general no guard addresses it
default-requirement/price_under studioneat.com — (the row renders no quote) Lowest readable price is $0.00. The listing is the merchant's own accidentally-published test record — title "SYDNEY TEST PRODUCT", sku equal to the variant id 49979497054485, OutOfStock, price 0 in JSON-LD, og:price:amount "0" and product.json "0.00". There is no product and no price, yet the merchant is told a price fact was evidenced. general no guard addresses it
default-requirement/price_under tenthousand.cc — (the row renders no quote) Lowest readable price is $0.00. The $0.00 is a masked/unset price on a real, purchasable garment, not a price. All 18 variants carry genuine SKUs (TTM0211301…) and GS1 barcodes (8404926790xx), several are available with live inventory (Iron/L shows 32 units, Iron/M 12, Salt/XL 7), yet price, price_min, price_max, og:price:amount and product.json are all 0 and the product is tagged exclude_include_in_collections. The merchant is told "Price under $10: pass" about a muscle tee that does not cost $0. general no guard addresses it

What this bound does NOT cover

  • ⚠️ SUPERSEDED IN PART, AND THE CORRECTION IS RECORDED RATHER THAN THE PARAGRAPH DELETED. The narrative below was written against the 18 confirmed defects of the v3.7 audit; this sample now records 11 over 483 rows, because v3.8 shipped two of the fixes it calls for. Mechanism (1) is CLOSED: the engine now refuses a price row outright when the store's own bytes declare a non-USD currency, rather than rendering a dollar sign over it. Mechanism (2) is CLOSED on the `/products/{handle}.json` tier, which now fails closed rather than passing an integer cents value through unchanged. Mechanisms (3) and (4) are OPEN and are filed as ENGINE_GAPS P-19. The paragraph is kept because its FINDING is what matters and is undiminished: the largest defect class in this engine was arithmetic rather than language, and no audit in this project's history had looked at it. Read what follows as the state that was measured, not as the state today. ⚠️ THE LARGEST DEFECT CLASS IN THIS ENGINE IS ARITHMETIC, NOT LANGUAGE, AND NO AUDIT IN THIS PROJECT'S HISTORY HAD LOOKED AT IT. 14 of the 18 are price rows. The kind looked like a tautology — the cap is generated by rounding the product's own price up, so the comparison always passes — and what there is to check is whether the NUMBER and the SENTENCE are true. Four mechanisms say no: (1) no code path in the engine reads a currency, so five stores publishing GBP/CAD/AUD/EUR/AUD are rendered with a US dollar sign while JSON-LD priceCurrency, Shopify.currency and Shopify.country all state otherwise on the same page; (2) priceToUsd's cents guard is `p > 1000 && Number.isInteger(p)`, so an integer CENTS value at or below 1000 from the /products/{handle}.js tier passes through unchanged — levainbakery.com's $10.00 mug is published as "Lowest readable price is $1000.00" on a strict `>` at the exact boundary, and richer-poorer.com's $3.00 item as $300.00; (3) $0.00 is treated as a price on five products that publish no price, including one whose own title is SYDNEY TEST PRODUCT and one Shopify `Referral` record; (4) on fieldcompany.com "Lowest readable price" is the page's MAXIMUM, because the JSON-LD publishes a single Offer at 135.00 while the analytics bootstrap on the same HTML lists a 7900-cent variant.
  • NONE OF THE 18 IS AN IDENTIFIER ROW. `identifiers` is 0 of 29. The floor this figure replaces was measured on that class alone and was right about it; what it was not is an error rate. A sample audited in one class tells you what that audit thought to look for.
  • ⚠️ THE PER-KIND CELLS DO NOT SUPPORT A SPREAD, and that is published rather than hidden behind a table. Of the 21 pairs of kinds in this sample, exactly one separates on non-overlapping 95% intervals, by a fraction of a percentage point. A six-cell table read as six rates would be false precision; the cells are published as COUNTS with intervals, and the pairwise test is published beside them.
  • Borderlines are counted as PASSES, the same convention as the coffee sample, and there are 13. Three are deliberate free gifts whose $0.00 the merchant's own title states, three are variant option values that are really the product's name, three are delivery timings whose scope is arguable, two are materials rows stating a property rather than a constituent, one is an ingredient-scoped organic claim, and one is a shipping sentence whose antecedent is the order rather than the product.
  • A general sample estimates the error rate on copy that looks like the average of every category at once, which is copy no individual merchant writes. That limit is unchanged by this measurement and is the reason a category standard is still fitness-measured on its own category.

What a passing claim row cannot rule out

What a passing claim row cannot rule out: on copy where the claim term sits next to a supplier, a farm, a region or a bundled item, the row establishes that the page states the term — not that the term was asserted of this product. Read the quoted sentence, which every passing row shows, and check what it is about.

We attack our own claim matcher with sentences written to break it. That measures CAPABILITY — what an adversary could do — and is deliberately not a measurement of what merchants write. Beside each capability figure is how often the same shape occurs in the sentences the engine actually rendered as evidence on real stores.

Attack shape, by capability and by real-copy frequency
What the sentence does Succeeds on sentences written to defeat it Occurs in real evidence sentences Status
The term is present, but the sentence does not assert it of this product — an invitation, a capability offer, a placeholder. 93.3% (252 of 270) 4.2% (3 of 71) not guarded — it does not occur
The property is described as past, future, conditional or merely possible rather than as holding now. 70.5% (425 of 603) 0% (0 of 71) not guarded — it does not occur
The term attaches to something other than the product — a supplier, a farm, a region, a bundled item, a practice. 48% (425 of 886) 15.5% (11 of 71) known limitation

The known limitation, and what it cost to try to close it

A guard for this axis was designed, implemented and measured in full. It closed 6 of its 8 real-copy targets and lost no true row on a 349-store replay. An independent adversarial pass — four attackers who wrote neither the guard nor its acceptance suite, 805 probes, every claimed regression re-executed by a refuter — confirmed 119 true statements it would have stopped reporting, against 6 defects closed. That is 19.8 true rows lost per defect closed, against a bar of 2.33–5.13. It was reverted and the limitation recorded as G-15.

The bar it had to clear was derived, not chosen: 14 honest carriers / 6 defects only this axis closes = 2.33; 41 / 8 = 5.13.

This table was completed on 2026-07-28 and deliberately NOT published until the open axis had either a fix or a recorded limitation — security-disclosure practice, decided before the outcome was known. Publishing selectively was considered and rejected: a table showing two axes at 'attacks well, occurs never' while omitting the third would read as a clean bill of health. It ships whole or not at all.

Measured 2026-07-28 against engine v2.4.0.

Measured 2026-07-28 against engine v3.5 rule D — unchanged at v3.6 and v3.7 (both measurement-only; no matcher file touched, asserted by gate).

Worked example: the row that shows you nothing

“A product identifier is published in structured data” is the one requirement in this standard whose result renders NO QUOTE. It reports that your structured data publishes an identifier; it cannot show you the identifier, so the row reads the same whether the value is a real barcode or a number the store made up about itself.

The row's promise is that a machine buyer can match this exact product to an EXTERNAL catalogue. A store-local id cannot do that — it resolves inside one shop and nowhere else — which is why the requirement names the SKU as explicitly insufficient.

3 real stores, from the captured bytes of their own pages. All 3 passed this row when the example was built; 2 still do. One of them always deserved to.

sputnikcoffeecompany.com — 600160850004

Honest pass: A 12-digit UPC-A whose GS1 mod-10 check digit validates (recomputed: 4, published: 4), published as gtin12 and mirrored into mpn and the variant barcode. This resolves outside the store.

The engine passes this row today. Your structured data publishes an MPN (600160850004).

…ts/8oz-sputnik-coffee-ground-vacuum-tin?variant=46026182066429" } ], "gtin12": "600160850004", "productId": "600160850004", "brand": { "name": "Sputnik Coffee Company"…

From the page captured on 2026-07-27, at byte 26656. Published in: gtin12, mpn, barcode.

glowrecipe.com — 8079462006899

Was a false pass; now refused: The same value appears in 11 different fields on this page, among them Shopify's own `productId` and `data-product-id`, and alongside `gid://shopify/Product/8079462006899`. It is the store's internal product id, written into `mpn` by the theme. The store's real SKU is a different string entirely.

The engine no longer passes this row. The only identifier we can use on your product's own structured-data node is an MPN (8079462006899), and that is the id your own storefront uses for this product — it resolves to nothing outside your store. We read that node only, so a GTIN published beneath it on an offer or a variant is not counted here.

…ca548 data-product-gallery-projects="[]" data-collection-gallery-projects="[]" data-product-id=8079462006899 data-template-name="product" data-ot-ignore data-cookieconsent="ignore" > </sc…

From the page captured on 2026-07-26, at byte 261002. Published in: rid, mpn, id, data-product-id, ProductID, resourceId, productId, data-prodID, data-productid, product_id, content_ids.

www.stumptowncoffee.com — 100754

False pass, still live: The same six digits are published as BOTH `sku` and `mpn`, adjacent in one JSON-LD object. A SKU is the store's own part number — the requirement's own insufficient-evidence list names it — so copying it into `mpn` cannot make it resolve anywhere else.

The engine passes this row today. Your structured data publishes an MPN (100754).

…d": "https:\/\/www.stumptowncoffee.com\/products\/net-wrecker", "@type": "Product","sku": "100754","mpn": "100754","brand": { "@type": "Brand", "name": "Stumptown Coffee" }, "description":…

From the page captured on 2026-07-27, at byte 1030. Published in: sku, mpn.

This is the class an audit cannot see. There is no rendered evidence to be suspicious of, so a reader checking every row learns nothing from any of them. The store-local values above sat inside the general sample below — 483 pass rows, every one read individually — and a reader checking each rendered quote would have found none of them. It took one mechanical check against the captured bytes. The bound this page publishes has moved three times as the audit method improved, and each move was a measurement of what the previous audit had thought to look for.

1 of these 2 is closed and 1 is still live, and the difference is the honest part. The engine now refuses an MPN that is byte-identical to the storefront's own product object id. It does NOT refuse one that is a copy of the store's own SKU: the rule that would (“an MPN equal to the SKU”) was scored over every MPN-publishing product in the corpus at 0 true positives and 7 false, and the seven are the compliant case — a brand that manufactures what it sells legitimately uses one string for both. Closing three sentences is not closing a class.

Record: experiments/v3-2/CATEGORY_BOUND.md; extracted mechanically by experiments/v3-3/extract_identifier_example.mjs · engine behaviour re-executed 2026-07-28 at 132d085

Measured discrimination, entry by entry

Each executable entry was run against the same recorded sample and its failure rate given a 95% interval, then compared with the 15-85% target band. THE INTERVAL DECIDES, not the point estimate — a rate outside the band whose interval straddles the edge has established nothing, and saying so is the difference between a measurement and a number.

Fail rate, 95% interval and verdict per entry
EntryFail rate95% interval Failed / adjudicatedVerdict
FORMAT-001 73.7% 64.3 – 81.4% 73 / 99 Discriminating
FORMAT-002 92.9% 86.1 – 96.5% 92 / 99 Not discriminating
GRIND-001 84.8% 76.5 – 90.6% 84 / 99 Discriminating
GRIND-002 92.9% 86.1 – 96.5% 92 / 99 Not discriminating
WEIGHT-001 49.0% 39.4 – 58.7% 49 / 100 Discriminating
CERT-001 92.0% 85.0 – 95.9% 92 / 100 Not discriminating
CERT-002 96.0% 90.2 – 98.4% 96 / 100 Not discriminating
SOURCE-001 89.0% 81.4 – 93.7% 89 / 100 Undecided
IDENT-001 97.4% 90.9 – 99.3% 74 / 76 Not discriminating
DELIV-001 60.8% 49.4 – 71.1% 45 / 74 (26 of 100 undecided by the engine, excluded) Discriminating

4 discriminating · 1 undecided · 5 not discriminating, of 10 measured entries.

Discriminating — the measured rate lies inside the target band, so the answer separates stores from one another.

Undecided — the rate is outside the band but its 95% interval is not — this measurement RAN AND DECIDED NOTHING, which is neither of the other two answers and must never be read as either. The entry stays executable and keeps accruing n.

Not discriminating — the whole 95% interval lies outside the target band, so almost every store answers the same way. This is evidence for a retirement decision and is not the decision.

1 of 10 decided nothing, and that is published rather than rounded to a decision. SOURCE-001 at 89.0% (95% interval 81.4–93.7%). An undecided result is "no difference detectable at this n", which is not "no difference".

4 of the 5 not-discriminating verdicts MAY NOT BE ACTED ON, and this document says so instead of working around it: FORMAT-002, GRIND-002, CERT-001, CERT-002. A confidence interval bounds sampling error and nothing else, and each of these was measured with a known bias in the instrument pointing at the very band edge its interval cleared. None of the entries below has been retired. A measured verdict is EVIDENCE for a retirement decision; it is not the decision.

Executable (10)

This question is asked as a test and produces a result for a real page.

Back to contents ↑

Not yet bound (5)

The engine can run this kind of check and a public product page can settle it — but THIS standard has not yet written the binding or put the entry through its adversarial pass. The obstacle is unwritten work in this document, not the evidence and not the engine. Each one names the engine capability that fits.

  • ALS-COFFEE-1.3-PRICE-001 — What does this cost? No engine ReqKind fits as written. `price_under` exists but adjudicates a price against a cap the buyer supplies, and this entry asks only whether a price is published at all; the engine exports no `price_is_stated` kind. No binding and no adversarial pass have been authored for this entry in this standard, so it is not `executable`; the obstacle is unwritten work in this document, not the evidence and not the engine. Because no kind fits, this one needs an engine-gap proposal rather than a binding invented here.
  • ALS-COFFEE-1.3-STOCK-001 — Can I actually buy this right now? `in_stock` is a live engine ReqKind and public product data adjudicates it, so this is neither `advisory` (the evidence exists) nor `blocked` (the engine exists). No binding and no adversarial pass have been authored for this entry in this standard, so it is not `executable`; the obstacle is unwritten work in this document, not the evidence and not the engine.
  • ALS-COFFEE-1.3-TERMS-001 — Can I buy a single bag without signing up for a recurring subscription? `no_subscription` is a live engine ReqKind and public product data adjudicates it. It is absence-based, so a conformant result is `pass_no_blocking` and never `pass_evidenced` — a binding here would have to say so explicitly. No binding and no adversarial pass have been authored for this entry in this standard, so it is not `executable`; the obstacle is unwritten work in this document, not the evidence and not the engine.
  • ALS-COFFEE-1.3-DECAF-004 — Can I buy a decaf version of this same coffee? No engine ReqKind fits: this entry adjudicates a statement about a DIFFERENT product from the one under test — whether a decaf counterpart exists elsewhere in the catalogue — and every kind the engine exports reads the product in front of it. No binding and no adversarial pass have been authored for this entry in this standard, so it is not `executable`; the obstacle is unwritten work in this document, not the evidence and not the engine. Because no kind fits, this one needs an engine-gap proposal rather than a binding invented here.
  • ALS-COFFEE-1.3-DIET-001 — Is this coffee vegan and gluten-free? `claim` is a live engine ReqKind and `vegan` and `gluten_free` are both live claim keys in the engine's own claim table, so public product copy adjudicates it. No binding and no adversarial pass have been authored for this entry in this standard, so it is not `executable`; the obstacle is unwritten work in this document, not the evidence and not the engine.

Back to contents ↑

Advisory (11)

This question is worth asking but is not reducible to a check a page can settle. It is published as guidance, not as a test.

Back to contents ↑

Blocked (16)

This question matters to a buyer and CANNOT be answered from a public product page today. What would be required is stated in full.

Back to contents ↑