01
The procedure is the output
Applies to marketplace onlyown storefronthybrid
This lesson publishes a procedure rather than findings. A result is assessable through the design that produced it, and a design stated after collection cannot be distinguished afterwards from one selected to fit what was found.
The source of a figure is part of the same question. Anyone selling software in this area, including the company publishing this course, has a commercial position on the answer, and the check available to a reader is a method specified in enough detail to be run independently. A figure arriving without one cannot be checked against anything.
02
What is fixed before the first query
Applies to marketplace onlyown storefronthybrid
The following are written into a dated document and not amended once results begin arriving.
- The exact prompts or queries — the wording rather than the theme — and how many times each will run.
- The scoring rubric, including what counts as the item appearing and what counts as a partial match.
- The sampling plan: which days, which times, which surfaces, which locale, and which account state.
- What counts as a negative result, in the same detail as a positive one.
- The limitations already known at the time of writing.
The fourth entry is the one the rest depends on. Where an unfavorable outcome has no written description in advance, the design retains flexibility in the direction the results later require, and that flexibility is not visible in the finished write-up.
03
Sampling a system that moves
Applies to marketplace onlyown storefronthybrid
These systems differ between days, between accounts, and between sessions, and they are personalized in ways a seller cannot fully suppress. A single observation therefore describes one session rather than a small quantity of the behavior.
| Recorded for every observation | What it accounts for |
|---|---|
| Surface, and any visible model or version identifier | A change here removes comparability across runs |
| Date and local time | Responses differ across days, and a rerun otherwise has no anchor |
| Locale and language | Availability and eligibility are frequently regional |
| Account state — signed in, signed out, history present | Personalization is recorded rather than removed |
| The verbatim response, not a summary of it | A summary is scoring, performed before the rubric is applied |
Observations are repeated across days rather than within one afternoon. Repetition inside a single session measures that session, which is the difference between the two designs.
04
Scoring blind, and reporting disagreement
Applies to marketplace onlyown storefronthybrid
Scoring happens after the collector knows what the run produced. Three steps separate the two: collection and scoring are performed as distinct passes, identifying context is stripped from the responses before scoring, and a second person scores a subset so the rate of disagreement between the two can be reported.
That disagreement rate is published rather than resolved privately. A rubric two people apply differently is a rubric one person applies differently across two weeks, and its size tells a reader how much of the reported figure is separable from the scorer.
The recommendation research cited here names factual consistency among the objectives its benchmark scores, which describes what that benchmark measures rather than a property of a deployed system.1 The transferable part is the sequence: the thing being scored is defined first, and scoring is performed against the definition.
05
First question: whether stock appears
Applies to marketplace onlyown storefronthybrid
Pointed at the visibility question, the procedure produces a bounded statement: for this prompt set, on these dates, on these surfaces, items of this kind appeared this often. That is a description of a sample, and it is recorded as one.
Two limits are properties of the design. It observes rather than intervenes, so it does not support a causal statement, and it does not extend past the prompts and dates sampled. Both are written into the result rather than appended to it.
The academic work in this area was conducted under laboratory conditions rather than against a storefront, and it does not describe how any commercial surface currently orders what it presents.2 A first-party observation with a stated method covers a different question from the one that literature asked.
06
Second question: how long a sold item stays buyable
Applies to marketplace onlyown storefronthybrid
The same procedure answers a second question: after a one-of-one item sells, how long does each surface continue to present it as purchasable? An appearance in an answer and a false availability value are different outcomes, and this measures the second.
- Sample real sold items across the channels and surfaces in use.
- Record the sale timestamp from the seller’s own records, to the minute.
- Check each surface on a fixed schedule and record the first check at which the item is no longer offered.
- Publish the distribution rather than the average alone, with the method and the sample size beside it.
| Surface | Observations | First check showing it gone | What the row supports |
|---|---|---|---|
| Own storefront | 6 items, 3 days | The first check, in every case | A statement about this shop |
| Marketplace listing | 6 items, 3 days | Within the same hourly check | The same — six items is a sample rather than a rate |
| A comparison site that had copied the page | 6 items, 3 days | Not observed — still offering two when sampling stopped | That the window exceeds the sampling period; not its length |
| A chat surface asked directly | 6 items, 3 days | Did not appear at all, before or after sale | A retrieval outcome; nothing about staleness |
The third row is left unfinished because that is what was observed. “Longer than the sampling period” is the statement the data carries, and an average would require values from outside it. The fourth row is filed under the first question rather than the second: an item that never appeared has no staleness window, and a zero there is not a staleness result.
One distinction separates the measurement from a product claim. The quantity here is the externally observable window on surfaces the seller does not control. A vendor’s internal synchronization latency is a different quantity, measured inside a system, and it is not what this procedure observes.
07
Publishing it in a re-runnable form
Applies to marketplace onlyown storefronthybrid
The pre-registration, the raw observations, the rubric, the disagreement rate, and the dates are published together. Raw data without the rubric and a headline without the data are each unusable for the same reason.
The decision to publish is made in advance and applies to both outcomes. A seller reporting that their items appeared in none of forty queries has published a result with a method attached; a favorable rate published without one is a figure that cannot be checked. Whether to publish is settled at the design stage, since settling it afterwards is settling it with the result in hand.
08
What a run of this size supports
Applies to marketplace onlyown storefronthybrid
The cited recommendation research evaluates factual consistency within a defined benchmark.1 The cold-start literature reports laboratory conditions rather than a live commercial surface.2
Neither publishes a merchant-side visibility rate or a staleness window, which is why a seller running this procedure is producing a primary observation. What it supports is a statement about the prompts, dates, and surfaces sampled, and the sample size is part of the statement rather than a caveat on it.
09
Practice
Exercise
Pre-register one small run
- Write ten exact prompts, a rubric, and your definition of a negative result, and date the document.
- Run them across three separate days, recording surface, time, locale, and account state.
- Score blind, have someone re-score a sample, and write the result with its limits attached.
Check yourself
Why is the rubric written before collection?
Because a rubric written afterwards is written by someone who has seen the results, and its flexibility is not visible in the finished write-up. Writing it first is what makes an unfavorable result reportable in the same form.
Why is one query on one day not a measurement?
These systems vary across days, accounts, and sessions, so a single observation records a session. Repetition across days is what separates the two.
Which staleness figure does this procedure produce?
The externally observable window during which a surface the seller does not control still offers a sold item. A vendor’s internal synchronization latency is measured inside a system and is a different quantity.
Progress is saved in this browser only. No account, nothing sent anywhere.
10
Common questions
How large does the sample need to be?
The size is reported alongside the result rather than being a threshold. A dozen items observed under a published method is a primary source; a hundred observed informally has a larger denominator and no method.
Can a tool that reports AI visibility be used instead?
Where it publishes its prompts, its sampling, and its scoring rule, the figure can be checked. Where it does not, the figure cannot be reproduced by the seller relying on it.
What if nothing appears anywhere?
That is a result with the same standing as any other. It closes a question that otherwise stays open, and the cited sources do not establish how common it currently is for one-of-one stock.
Is it worth measuring something the seller cannot change?
The staleness measurement points at a value the seller sets directly. The visibility measurement sizes a question, and its repetition cadence is a choice recorded with the design.
11
Research and sources
Rules and platform policies change. These primary sources were reviewed on ; confirm the current position for your jurisdiction and account before acting.
Claim evidence
- Recommendation research treats factual consistency as an objective evaluated inside a defined benchmark rather than as a property assertable about a live commercial system.
- observed outcome. Supported by Factual and Personalized Recommendation Language Modeling with Reinforcement Learning .
- Evaluations of language-model recommenders were run in research settings that are not merchant feeds and do not expose any platform’s live ranking behavior.
- observed outcome. Supported by Large Language Models are Competitive Near Cold-start Recommenders for Language- and Item-based Preferences .
- Factual and Personalized Recommendation Language Modeling with Reinforcement Learning Google · Tier B · peer reviewed · COLM 2024 proceedings
Factual consistency and personalization in the paper’s conversational-recommendation benchmark. Limit: The MovieLens-based study provides design evidence, not proof of current shopping-platform behavior.
- Large Language Models are Competitive Near Cold-start Recommenders for Language- and Item-based Preferences Google · Tier B · peer reviewed · RecSys 2023 proceedings
Language-based preference recommendations in the paper’s near-cold-start experiment. Limit: The evaluated recommender setting is not a merchant feed and does not expose ChatGPT ranking behavior.