Last updated 3 October 2026. The current AI commerce reliability project is an exploratory pilot. Its checks are AI-assisted and provisional; independent human review is pending.
Exploratory and pre-registered work
Each project states its research question, design, evidence status and review status. Exploratory work can develop a question or method without being presented as a confirmatory test.
A pre-registration records a dated question, design and analysis plan before the relevant evidence is examined. Publishing a protocol after looking at results does not make those results pre-registered. An exploratory pilot may inform a later test using fresh or genuinely held-out evidence. The pilot keeps its exploratory label.
Research updates and methods explainers are labelled by their purpose. They do not become confirmatory research simply because they are published here.
Amendments
Changes to a research question, prompt set, rubric or analysis are recorded with a date, a reason and the affected version. Changes are made before collecting or analysing the observations to which the revised rule will apply.
Original observations are preserved. A new rule does not silently rewrite what was collected under an earlier design. The project page links to the relevant protocol and change history when those materials are publicly released.
Review
Review credits state who checked the work and what they checked: the code, evidence, rubric, annotations or conclusions. Those roles are distinct, and a review of one part does not imply approval of everything.
AI assistance is disclosed. Software completeness checks and AI-assisted judgments do not establish that an annotation is true and do not count as independent human validation.
For the current pilot, independent human review has not been completed. Its protocol calls for a second person to review at least four complete answers across surfaces and difficult labels before the pilot’s results are presented as validated.
Corrections
Errors in an explanation or annotation are corrected with a dated record of what changed and why. Raw responses remain unchanged. A correction preserves the distinction between the original observation and a revised interpretation.
Material changes to a conclusion are noted on the project page and linked to the underlying correction history. The repository change log, once released, is the authoritative record; the project page carries a dated summary.
Disagreements
Substantive methodological objections and unresolved reviewer disagreements are documented alongside the relevant work. Responses explain the evidence and reasoning rather than treating agreement as a condition for publication.
Outside contributors and reviewers are credited for their actual contribution. Private correspondence is attributed only with permission.
Simulation
Simulations are used where they fit the question, to examine how a method behaves when the generating process is known. Their assumptions and sensitivity to alternative assumptions are stated.
A result from simulated commerce data is a result about that simulation. It does not establish what happened in a real business. Simulated observations are clearly labelled and kept separate from real observations.
Data
Current research uses public, purpose-collected, synthetic or consented data. It uses no data from any current or former employer.
Data records identify their source, collection timing, scope, missingness and limitations. Public sharing respects the applicable consent and rights. A project’s materials are explicitly marked as available, withheld or not yet released.
Earlier practitioner articles are retained in Measurement Guides and CRO Lab. They were written before these standards and are not presented as research validated under them.
What a completed study includes
A completed study states its question, method, findings and limitations, with review credits and a section titled “When this is wrong.” It links to its registration where applicable and to publicly released code, evidence and data where available. Missing materials or unfinished checks are stated explicitly.The first project, AI shopping reliability: the backpack pilot, asks whether AI shopping recommendations meet shoppers’ requirements and can be judged reliably.

