# https://lens.thirdocular.com > Versioned buying standards: the questions a competent buyer asks in a category, > written as executable tests and run against a store's real product pages. > Each requirement reports proven, not proven, or requires store access, with the > evidence that decided it. We publish what we cannot test, and why. ## Standards - [AisleLens Coffee Standard v1.0 — buyer questions, assertions and evidence rules for roasted coffee product pages](https://lens.thirdocular.com/standards/coffee/1.0): 42 entries (10 executable, 5 not_discriminating, 11 advisory, 16 blocked). Grammar 1.0. Content hash 334389c4eb6145112deec621e667f11142fb204c66bedd314fc12662d09acec5. - SUPERSEDED by v1.1 (https://lens.thirdocular.com/standards/coffee/1.1). Served unchanged so existing citations resolve; do not cite it for new work. - [JSON](https://lens.thirdocular.com/standards/coffee/1.0/standard.json) — the artifact a citation resolves against. - [Grounding](https://lens.thirdocular.com/standards/coffee/1.0/grounding) — every external source, and what it establishes. - Measured false-positive rate, Coffee category sample: 6.17% point estimate, 12.78% 95% upper bound (n=162 pass rows across 77 stores, cluster-adjusted). - Measured false-positive rate, General DTC sample: 3.54% point estimate, 7.80% 95% upper bound (n=509 pass rows across 169 stores, cluster-adjusted) — a FLOOR: only one defect class was re-checked, so the true rate is at least this. - [AisleLens Coffee Standard v1.1 — buyer questions, assertions and evidence rules for roasted coffee product pages](https://lens.thirdocular.com/standards/coffee/1.1): 42 entries (10 executable, 5 not_discriminating, 11 advisory, 16 blocked). Grammar 1.1. Content hash f8ec2780f60c38931913e5b6cd37506500c8462709209de7180ba6691d6137e7. - SUPERSEDED by v1.2 (https://lens.thirdocular.com/standards/coffee/1.2). Served unchanged so existing citations resolve; do not cite it for new work. - [JSON](https://lens.thirdocular.com/standards/coffee/1.1/standard.json) — the artifact a citation resolves against. - [Grounding](https://lens.thirdocular.com/standards/coffee/1.1/grounding) — every external source, and what it establishes. - Measured false-positive rate, Coffee category sample: 6.17% point estimate, 12.78% 95% upper bound (n=162 pass rows across 77 stores, cluster-adjusted). - Measured false-positive rate, General DTC sample: 3.54% point estimate, 7.80% 95% upper bound (n=509 pass rows across 169 stores, cluster-adjusted) — a FLOOR: only one defect class was re-checked, so the true rate is at least this. - [AisleLens Coffee Standard v1.2 — buyer questions, assertions and evidence rules for roasted coffee product pages](https://lens.thirdocular.com/standards/coffee/1.2): 42 entries (10 executable, 5 unbound, 11 advisory, 16 blocked). Grammar 1.2. Content hash fe199a864d3d4d565986851f9bfae9e108d55e4c86af18b1f8027f3d23486b58. - SUPERSEDED by v1.3 (https://lens.thirdocular.com/standards/coffee/1.3). Served unchanged so existing citations resolve; do not cite it for new work. - [JSON](https://lens.thirdocular.com/standards/coffee/1.2/standard.json) — the artifact a citation resolves against. - [Grounding](https://lens.thirdocular.com/standards/coffee/1.2/grounding) — every external source, and what it establishes. - Measured false-positive rate, ALS-COFFEE category sample: 6.17% point estimate, 12.78% 95% upper bound (n=162 pass rows across 77 stores, cluster-adjusted). - Completion state: DEFECTS_FOUND. 10 confirmed false passes. A measurement that did not finish resolves to INCOMPLETE and may never be summed into a defect total or read as a pass — zero is the most dangerous number a broken instrument returns, because it is also what a healthy one returns. - Limit: This bounds SAMPLING error and nothing else. It does not cover a bias in the engine doing the measuring, and this run has one: the bounded semantic tier was disabled. Both of its paths are therefore absent — the missing GRANT path pushes a claim row's fail rate up and the missing VETO path pushes it down — and nothing here measures which dominates, so the tier is declared as two biases per affected entry rather than netted into one. - Limit: One product per storefront. 100 products over 100 distinct hosts means within-store copy variation is entirely unmeasured: nothing here says a second product from the same roaster would fare the same. - Limit: The comparison with the general DTC sample is a FLOOR, not a bound, and the two are NOT audited to the same depth. The general figure — 7.8% cluster-adjusted over 509 pass rows from 169 stores — rests on ONE defect class re-checked mechanically, after an audit that read every rendered row and confirmed zero. So the ratio between 7.8% and 12.78% is not itself a measurement; the DIRECTION is what the category rule rests on. - Limit: 10 confirmed false positives is a count of rows an auditor could recognise as wrong from the page text. A defect class that renders no quote is invisible to any audit that reads rendered evidence, however many rows it reads — which is how a previously published 0.83% general bound turned out to be 7.80%. - Limit: 7 rows were judgement calls counted as PASSES. Counting them as defects instead would raise every figure here; they are named in the audit record so the call can be disputed rather than merely trusted. - Limit: Fitness was measured on the ten entries that are executable. It says nothing about the 32 entries that are not run, including the 5 now marked `unbound`. - Measured discrimination over 10 executable entries: 4 discriminating, 1 indeterminate, 5 not discriminating. `indeterminate` means the measurement ran and decided nothing; it is not either of the other two. None has been retired. - [AisleLens Coffee Standard v1.3 — buyer questions, assertions and evidence rules for roasted coffee product pages](https://lens.thirdocular.com/standards/coffee/1.3): 42 entries (10 executable, 5 unbound, 11 advisory, 16 blocked). Grammar 1.2. Content hash ba2050578ed0274885fd6213967c230b2a57dd2b7c1d3fba8c5e1633027d4cf7. - CURRENT version of coffee. Cite this one. - [JSON](https://lens.thirdocular.com/standards/coffee/1.3/standard.json) — the artifact a citation resolves against. - [Grounding](https://lens.thirdocular.com/standards/coffee/1.3/grounding) — every external source, and what it establishes. - Measured false-positive rate, Coffee category sample: 4.38% point estimate, 9.99% 95% upper bound (n=160 pass rows across 77 stores, cluster-adjusted). - Completion state: DEFECTS_FOUND. 7 confirmed false passes. A measurement that did not finish resolves to INCOMPLETE and may never be summed into a defect total or read as a pass — zero is the most dangerous number a broken instrument returns, because it is also what a healthy one returns. - Limit: ⚠️ THE STUMPTOWN ROW IS RE-SCORED AS A TRUE PASS, AND THIS SUPERSEDES THIS SIDECAR'S OWN EARLIER READING. www.stumptowncoffee.com publishes "sku":"100754","mpn":"100754". The earlier reading counted it as a surviving defect the engine deliberately does not catch. Scored against ALS-COFFEE v1.3's ACTUAL TEXT it is not a defect at all, so there is nothing for the engine to catch and counting it inflated the bound. The deciding sentence is IDENT-001's own adversarial.residual_risk clause (2), verbatim: "It disqualifies a value for being the seller's own object id, not for being seller-private in general: a stock code that is neither a placeholder nor the storefront's key is outside this clause." accepted_evidence and insufficient_evidence scope identically. Checked against the bytes rather than against that reasoning: the storefront's own product key is 9516469289128 and its variant key is 55754751967400 (engine shopifyStorefrontObjectId and meta.product.id agree); 100754 is the merchant's SKU. The third key carrying the value is the analytics defaultVariant.id / items[].id, whose conventional value in a GA-shaped payload IS the SKU — reading that as "the storefront's key" would disqualify every published SKU, which is the reading residual_risk (2) forecloses and the rule (mpn === sku) already measured at 0 true positives and 7 false. - Limit: ⚠️ A TENSION INSIDE THE ENTRY, STATED RATHER THAN RESOLVED SILENTLY. IDENT-001's insufficient_evidence clause for the MPN field reasons field-agnostically in its why_not — "A seller-private stock code is not a global identifier ... That reason is field-agnostic" — which read alone would disqualify a SKU echo. Its `form` and its residual_risk do not. `form` is the operative text an evaluator matches against and why_not is justification prose, so the rule governs; but the document owes a narrower why_not or a wider form, and until it has one this row is decided by a reading rather than by a match. - Limit: The bound covers the engine's false PASSES only. False fails are not in it, and rule D creates them: six real merchants in the captured corpus publish a check-digit-valid GTIN on a Product node the extractor does not select, and are now told they publish no usable identifier (ENGINE_GAPS P-12). - Limit: The comparison with the general DTC sample is now like for like for the first time — both samples have had every passing row adjudicated individually — and the two are STATISTICALLY INDISTINGUISHABLE. See cross_sample_comparison. No ratio is stated between them, because their intervals overlap. - Measured false-positive rate, General DTC sample: 2.28% point estimate, 5.17% 95% upper bound (n=483 pass rows across 169 stores, cluster-adjusted). - Completion state: DEFECTS_FOUND. 11 confirmed false passes. A measurement that did not finish resolves to INCOMPLETE and may never be summed into a defect total or read as a pass — zero is the most dangerous number a broken instrument returns, because it is also what a healthy one returns. - Limit: ⚠️ SUPERSEDED IN PART, AND THE CORRECTION IS RECORDED RATHER THAN THE PARAGRAPH DELETED. The narrative below was written against the 18 confirmed defects of the v3.7 audit; this sample now records 11 over 483 rows, because v3.8 shipped two of the fixes it calls for. Mechanism (1) is CLOSED: the engine now refuses a price row outright when the store's own bytes declare a non-USD currency, rather than rendering a dollar sign over it. Mechanism (2) is CLOSED on the `/products/{handle}.json` tier, which now fails closed rather than passing an integer cents value through unchanged. Mechanisms (3) and (4) are OPEN and are filed as ENGINE_GAPS P-19. The paragraph is kept because its FINDING is what matters and is undiminished: the largest defect class in this engine was arithmetic rather than language, and no audit in this project's history had looked at it. Read what follows as the state that was measured, not as the state today. ⚠️ THE LARGEST DEFECT CLASS IN THIS ENGINE IS ARITHMETIC, NOT LANGUAGE, AND NO AUDIT IN THIS PROJECT'S HISTORY HAD LOOKED AT IT. 14 of the 18 are price rows. The kind looked like a tautology — the cap is generated by rounding the product's own price up, so the comparison always passes — and what there is to check is whether the NUMBER and the SENTENCE are true. Four mechanisms say no: (1) no code path in the engine reads a currency, so five stores publishing GBP/CAD/AUD/EUR/AUD are rendered with a US dollar sign while JSON-LD priceCurrency, Shopify.currency and Shopify.country all state otherwise on the same page; (2) priceToUsd's cents guard is `p > 1000 && Number.isInteger(p)`, so an integer CENTS value at or below 1000 from the /products/{handle}.js tier passes through unchanged — levainbakery.com's $10.00 mug is published as "Lowest readable price is $1000.00" on a strict `>` at the exact boundary, and richer-poorer.com's $3.00 item as $300.00; (3) $0.00 is treated as a price on five products that publish no price, including one whose own title is SYDNEY TEST PRODUCT and one Shopify `Referral` record; (4) on fieldcompany.com "Lowest readable price" is the page's MAXIMUM, because the JSON-LD publishes a single Offer at 135.00 while the analytics bootstrap on the same HTML lists a 7900-cent variant. - Limit: NONE OF THE 18 IS AN IDENTIFIER ROW. `identifiers` is 0 of 29. The floor this figure replaces was measured on that class alone and was right about it; what it was not is an error rate. A sample audited in one class tells you what that audit thought to look for. - Limit: ⚠️ THE PER-KIND CELLS DO NOT SUPPORT A SPREAD, and that is published rather than hidden behind a table. Of the 21 pairs of kinds in this sample, exactly one separates on non-overlapping 95% intervals, by a fraction of a percentage point. A six-cell table read as six rates would be false precision; the cells are published as COUNTS with intervals, and the pairwise test is published beside them. - Limit: Borderlines are counted as PASSES, the same convention as the coffee sample, and there are 13. Three are deliberate free gifts whose $0.00 the merchant's own title states, three are variant option values that are really the product's name, three are delivery timings whose scope is arguable, two are materials rows stating a property rather than a constituent, one is an ingredient-scoped organic claim, and one is a shipping sentence whose antecedent is the order rather than the product. - Limit: A general sample estimates the error rate on copy that looks like the average of every category at once, which is copy no individual merchant writes. That limit is unchanged by this measurement and is the reason a category standard is still fitness-measured on its own category. - Measured discrimination over 10 executable entries: 4 discriminating, 1 indeterminate, 5 not discriminating. `indeterminate` means the measurement ran and decided nothing; it is not either of the other two. None has been retired. ## A worked result - [An Example test on a real store](https://lens.thirdocular.com/demo) — a published standard executed against a real coffee product page, every requirement with the evidence sentence and the surface it was read from. ## Method - [Methodology](https://lens.thirdocular.com/methodology)