Wednesday Sep 2

Penn Finds AI Shoppers Unpredictable

2SEP
SAME ASKNEW PICK

AI shopping agents are not stable buyers. Penn researchers ran 200 trials across six frontier models. Adding one page of prior content, or just reordering what the agent read, changed which product it chose.

With no extra context, every model had one favourite item and stuck to it. Context broke that.

Direction of the swing depends on the model, which sources turn up, their order, and how results are bundled into tool calls. A seller sees none of those.

Sometimes the winner was worse on price, rating and review count than the item it beat.

full brief & sources

⚡ Why this matters

  • Agentic commerce is being built on the assumption that a good product wins. This says the retrieval path wins.
  • For anyone selling online, this is worse than SEO. SEO had feedback. Here you cannot see the model, the harness, or what it already read.
  • It is also a general warning about agent evaluation: single-shot benchmarks understate how much real deployments wobble.

🔍 What happened

  • New working paper led by University of Pennsylvania researchers, including Ethan Mollick.
  • 200 runs per condition across Claude Haiku 4.5 and Opus 4.8, GPT-5 Mini and GPT-5.5, Gemini 3.1 Flash Lite and Gemini 3.5 Flash.
  • Agents were shown reviews, recommendations, search results and user memories before choosing.
  • Adding prior content, changing its order, or repackaging it into different tool calls all moved the pick.
  • Authors: 'Two users issuing the same request, or the same user on a different day or a different model, may receive different products without any visible explanation.'
  • Their read for sellers: 'limited control rather than new leverage'.
  • They recommend robustness testing with adversarial prior content, and flagging influential user-memory statements.

💬 Smart takes

  • The authors' sharpest line is that agents have no mechanism to discount planted prior content, unlike a human who can be told an ad is an ad.
  • Forbes frames it as a trust problem, arriving just as surveys show most shoppers already act on AI recommendations without checking.
  • The uncomfortable corollary: whoever controls the retrieval harness controls the purchase, not whoever makes the product.

🧭 Where this goes

  1. Likelyan 'agent robustness' line item shows up in ecommerce vendor pitches within two quarters.
  2. Likelya cottage industry selling prior-content placement for agents, sold as GEO.
  3. Possiblea platform ships provenance flags on retrieved content specifically to stabilise agent purchases.
  4. Wild Carda regulator treats planted prior content aimed at agents as deceptive advertising.

🥄 The Spoon Take

This is the study to hand anyone who says agentic commerce is nearly solved. The models are not choosing badly, they are choosing unstably, and instability is harder to fix than bias. If your product roadmap assumes an agent will reliably find the better option, that assumption now has a number attached to it.

🤔 Pushback

It is a simulated shopping task in a working paper, not live checkout behaviour, and 200 runs per condition is small for claims about six different models.