Skip to main content

Part 06 · Guide 26 of 33

A method for measuring what agents say about your stock

This lesson covers why the procedure rather than the finding is the output, what is fixed in writing before the first query, how to sample a system that changes between days and accounts, how to score without the scorer knowing what they are looking at, the two questions this procedure answers for a reseller, and what publishing it requires.

Reading time
13 min
Sections
08
Last updated
July 30, 2026
Published by
Instica

01

The procedure is the output

Applies to marketplace onlyown storefronthybrid

This lesson publishes a procedure rather than findings. A result is assessable through the design that produced it, and a design stated after collection cannot be distinguished afterwards from one selected to fit what was found.

The source of a figure is part of the same question. Anyone selling software in this area, including the company publishing this course, has a commercial position on the answer, and the check available to a reader is a method specified in enough detail to be run independently. A figure arriving without one cannot be checked against anything.

02

What is fixed before the first query

Applies to marketplace onlyown storefronthybrid

The following are written into a dated document and not amended once results begin arriving.

  1. The exact prompts or queries — the wording rather than the theme — and how many times each will run.
  2. The scoring rubric, including what counts as the item appearing and what counts as a partial match.
  3. The sampling plan: which days, which times, which surfaces, which locale, and which account state.
  4. What counts as a negative result, in the same detail as a positive one.
  5. The limitations already known at the time of writing.

The fourth entry is the one the rest depends on. Where an unfavorable outcome has no written description in advance, the design retains flexibility in the direction the results later require, and that flexibility is not visible in the finished write-up.

03

Sampling a system that moves

Applies to marketplace onlyown storefronthybrid

These systems differ between days, between accounts, and between sessions, and they are personalized in ways a seller cannot fully suppress. A single observation therefore describes one session rather than a small quantity of the behavior.

Recorded for every observationWhat it accounts for
Surface, and any visible model or version identifierA change here removes comparability across runs
Date and local timeResponses differ across days, and a rerun otherwise has no anchor
Locale and languageAvailability and eligibility are frequently regional
Account state — signed in, signed out, history presentPersonalization is recorded rather than removed
The verbatim response, not a summary of itA summary is scoring, performed before the rubric is applied

Observations are repeated across days rather than within one afternoon. Repetition inside a single session measures that session, which is the difference between the two designs.

04

Scoring blind, and reporting disagreement

Applies to marketplace onlyown storefronthybrid

Scoring happens after the collector knows what the run produced. Three steps separate the two: collection and scoring are performed as distinct passes, identifying context is stripped from the responses before scoring, and a second person scores a subset so the rate of disagreement between the two can be reported.

That disagreement rate is published rather than resolved privately. A rubric two people apply differently is a rubric one person applies differently across two weeks, and its size tells a reader how much of the reported figure is separable from the scorer.

The recommendation research cited here names factual consistency among the objectives its benchmark scores, which describes what that benchmark measures rather than a property of a deployed system.1 The transferable part is the sequence: the thing being scored is defined first, and scoring is performed against the definition.

05

First question: whether stock appears

Applies to marketplace onlyown storefronthybrid

Pointed at the visibility question, the procedure produces a bounded statement: for this prompt set, on these dates, on these surfaces, items of this kind appeared this often. That is a description of a sample, and it is recorded as one.

Two limits are properties of the design. It observes rather than intervenes, so it does not support a causal statement, and it does not extend past the prompts and dates sampled. Both are written into the result rather than appended to it.

The academic work in this area was conducted under laboratory conditions rather than against a storefront, and it does not describe how any commercial surface currently orders what it presents.2 A first-party observation with a stated method covers a different question from the one that literature asked.

06

Second question: how long a sold item stays buyable

Applies to marketplace onlyown storefronthybrid

The same procedure answers a second question: after a one-of-one item sells, how long does each surface continue to present it as purchasable? An appearance in an answer and a false availability value are different outcomes, and this measures the second.

  1. Sample real sold items across the channels and surfaces in use.
  2. Record the sale timestamp from the seller’s own records, to the minute.
  3. Check each surface on a fixed schedule and record the first check at which the item is no longer offered.
  4. Publish the distribution rather than the average alone, with the method and the sample size beside it.
SurfaceObservationsFirst check showing it goneWhat the row supports
Own storefront6 items, 3 daysThe first check, in every caseA statement about this shop
Marketplace listing6 items, 3 daysWithin the same hourly checkThe same — six items is a sample rather than a rate
A comparison site that had copied the page6 items, 3 daysNot observed — still offering two when sampling stoppedThat the window exceeds the sampling period; not its length
A chat surface asked directly6 items, 3 daysDid not appear at all, before or after saleA retrieval outcome; nothing about staleness

The third row is left unfinished because that is what was observed. “Longer than the sampling period” is the statement the data carries, and an average would require values from outside it. The fourth row is filed under the first question rather than the second: an item that never appeared has no staleness window, and a zero there is not a staleness result.

One distinction separates the measurement from a product claim. The quantity here is the externally observable window on surfaces the seller does not control. A vendor’s internal synchronization latency is a different quantity, measured inside a system, and it is not what this procedure observes.

07

Publishing it in a re-runnable form

Applies to marketplace onlyown storefronthybrid

The pre-registration, the raw observations, the rubric, the disagreement rate, and the dates are published together. Raw data without the rubric and a headline without the data are each unusable for the same reason.

The decision to publish is made in advance and applies to both outcomes. A seller reporting that their items appeared in none of forty queries has published a result with a method attached; a favorable rate published without one is a figure that cannot be checked. Whether to publish is settled at the design stage, since settling it afterwards is settling it with the result in hand.

08

What a run of this size supports

Applies to marketplace onlyown storefronthybrid

The cited recommendation research evaluates factual consistency within a defined benchmark.1 The cold-start literature reports laboratory conditions rather than a live commercial surface.2

Neither publishes a merchant-side visibility rate or a staleness window, which is why a seller running this procedure is producing a primary observation. What it supports is a statement about the prompts, dates, and surfaces sampled, and the sample size is part of the statement rather than a caveat on it.

09

Practice

Exercise

Pre-register one small run

  1. Write ten exact prompts, a rubric, and your definition of a negative result, and date the document.
  2. Run them across three separate days, recording surface, time, locale, and account state.
  3. Score blind, have someone re-score a sample, and write the result with its limits attached.

Check yourself

Why is the rubric written before collection?

Because a rubric written afterwards is written by someone who has seen the results, and its flexibility is not visible in the finished write-up. Writing it first is what makes an unfavorable result reportable in the same form.

Why is one query on one day not a measurement?

These systems vary across days, accounts, and sessions, so a single observation records a session. Repetition across days is what separates the two.

Which staleness figure does this procedure produce?

The externally observable window during which a surface the seller does not control still offers a sold item. A vendor’s internal synchronization latency is measured inside a system and is a different quantity.

10

Common questions

How large does the sample need to be?

The size is reported alongside the result rather than being a threshold. A dozen items observed under a published method is a primary source; a hundred observed informally has a larger denominator and no method.

Can a tool that reports AI visibility be used instead?

Where it publishes its prompts, its sampling, and its scoring rule, the figure can be checked. Where it does not, the figure cannot be reproduced by the seller relying on it.

What if nothing appears anywhere?

That is a result with the same standing as any other. It closes a question that otherwise stays open, and the cited sources do not establish how common it currently is for one-of-one stock.

Is it worth measuring something the seller cannot change?

The staleness measurement points at a value the seller sets directly. The visibility measurement sizes a question, and its repetition cadence is a choice recorded with the design.

11

Research and sources

Rules and platform policies change. These primary sources were reviewed on ; confirm the current position for your jurisdiction and account before acting.

Claim evidence

Recommendation research treats factual consistency as an objective evaluated inside a defined benchmark rather than as a property assertable about a live commercial system.
observed outcome. Supported by Factual and Personalized Recommendation Language Modeling with Reinforcement Learning .
Evaluations of language-model recommenders were run in research settings that are not merchant feeds and do not expose any platform’s live ranking behavior.
observed outcome. Supported by Large Language Models are Competitive Near Cold-start Recommenders for Language- and Item-based Preferences .

Start with 25 items. Stay for 25,000.

Free for 25 items · No card · Cancel from your account page.

También disponible en españolEspañol →
Disponível em portuguêsPortuguês →
Auf Deutsch verfügbarDeutsch →
Disponible en françaisFrançais →
中文版本可用中文 →