Exploratory pilot · not pre-registered · AI-assisted checks; independent human review pending · page updated 3 October 2026
Observations described here were collected on 28 September 2026. Updating this page does not mean new observations have been collected.
The question
Do AI shopping recommendations satisfy shoppers’ stated requirements, and can that be judged reliably?
I am starting with backpacks sold online in India. Before counting errors or comparing systems, I want to establish whether a reviewer can identify the exact recommended product, find relevant evidence and make a defensible judgment.
Why it matters
A recommendation can name a product while leaving a shopper unable to confirm its specifications, price or purchase route. For a commerce team, the practical question is which information problems can be verified and which might justify a correction.
The pilot will inform a first decision: whether this review method is feasible enough to extend, needs a narrower question, or should change direction. It does not yet establish demand for an audit service, the effectiveness of a merchant correction, or an effect on sales.
Current scope
The planned design has five researcher-written English prompts, two consumer search surfaces and two runs per prompt: 20 answers in total. The surfaces are ChatGPT Search and Perplexity Search. These are buying tasks designed for a feasibility pilot, not a representative sample of Indian shoppers.
Each prompt asks for up to three specific products, a current listed price and a purchase link. Conditional bank or membership discounts are excluded. The five requirement families are:
Laptop fit: below ₹4,000 and compatible with a 15.6-inch laptop.
Capacity: below ₹3,000 and a stated capacity of at least 25 litres.
Weight: below ₹4,000 and manufacturer-published product weight no more than 800 grams.
Warranty: below ₹5,000 and a manufacturer warranty of at least one year applicable in India.
Dimensions: below ₹4,000 and manufacturer-listed external dimensions no greater than 45 cm high, 35 cm wide and 20 cm deep. This does not establish airline approval.
Every prompt also requires availability for online purchase in India. Availability is checked for the recommended merchant; it does not establish delivery to every postcode.
Currently, three Perplexity answers have been saved, containing nine product recommendations. Independent human review has not been completed for any answer. One further prompt was submitted but its answer could not be recovered; another has not been submitted. The planned ChatGPT Search observations have not been collected.
What exists
The working materials include:
A versioned protocol and five fixed prompts.
A requirement rubric separating supported constraints, constraint violations, insufficient evidence and pending review.
Saved answer transcripts, collection records and dated verification excerpts.
Provisional review records, including unresolved judgments and correction history.
Code for capturing answers, protecting records from accidental overwrite, checking review completeness and generating descriptive reports.
The saved answers are transcripts from consumer interfaces, not native exports or API responses. Records include the surface, displayed mode or model label and collection timing. An exact underlying model version is not assumed when the interface does not disclose it.
Public repository links: pending. The working materials exist locally, but this page does not claim they have already been publicly released. Reviewer-agreement calculations have not yet been implemented or completed.
What we have learned so far
These are preliminary lessons about the review process, not validated estimates of platform accuracy:
Exact product and variant matching matters before a specification can support a judgment.
A price or stock page checked after an answer establishes its state at verification. It may not establish its state when the answer was generated.
Missing evidence does not demonstrate that a product claim is false. An unverifiable claim needs a different label from an established violation.
Collection and source-access failures must remain visible. They must not be turned into apparent recommendation errors or silently replaced with different interfaces.
The current records are insufficient for an accuracy estimate, platform ranking or causal claim. The original provisional annotations also include unresolved timing interpretations and are not pooled into a headline result.
Open questions
Can another person apply the requirement rubric consistently? Which ambiguous cases need clearer rules? How much active verification effort does each answer require? Would a commerce practitioner identify a plausible action from the evidence?
A later question is how often a brand appears in sampled answers, and how many prompts, runs and days would be needed to detect a meaningful change. That is a separate measurement problem. Visibility alone would not establish recommendation quality, customer reach or incremental sales.
When this is wrong
The conclusions would be too strong if they treated these five prompts as representative of shoppers, counted repeated answers or product checks as independent people, or mistook later stock information for generation-time truth.
Manufacturer statements establish published specifications, not independently tested product performance. Merchant offers do not establish nationwide availability. The prompts request sources and explicit uncertainty, so findings apply to that wording.
The current review is AI-assisted and provisional. Independent human review, reviewer agreement, generation-time stock and price, some exact variant matches, and per-answer verification effort remain unresolved. Account context, location and underlying model versions are incompletely known. None of this pilot establishes the cause of an answer, the effect of changing a merchant page, or business value.
Change log
28 September 2026: protocol v0.1 and the working capture/review process were documented. The pilot remains exploratory and is not registered as a confirmatory study.
2 October 2026: this project introduction was prepared from the existing records.
3 October 2026: the project page was published. No additional observations or completed human reviews were added.
The next checkpoint is a five-answer evidence table and a decision about whether to continue, narrow the question or change direction. Expansion depends on that checkpoint.
Background
This project follows earlier work on the limits of AI-assisted conversion optimisation: I Built an AI CRO System. Here’s Where It Failed. That article is retained as background, separate from the current pilot’s evidence.
For the lab’s documentation and review policies, see Research Standards.

