Researchers at the Wharton School at the University of Pennsylvania tested how consistently AI shopping agents recommend products when the search process changes. Even tiny shifts in context swung purchase decisions by wide margins.
The team tested six current models, both mini variants and frontier-level, tasking each one to act as a personal shopping assistant picking a fitness watch from a fixed product grid. They used the ACES simulator (Agentic e-Commerce Simulator), which shows the AI agent a screenshot of a product page. The agent analyzes the image, optionally pulls in recommendation sources, and then picks a product.

A single source is enough to flip the recommendation
Even without external sources, the models showed different baseline preferences. But when the agent saw just one external source before the product page, recommendations shifted dramatically in some cases.
The researchers tested three sources: a Reddit thread recommending the Garmin Forerunner 55, a Wirecutter review for the Fitbit Inspire 3, and a Strategist article about the WHOOP 5.0. Wirecutter had the strongest pull. The probability of picking the Fitbit Inspire 3 jumped by 90 percentage points for Claude Opus 4.8 compared to the control condition, and by 99 percentage points for Gemini 3.5 Flash.

In a second experiment, agents saw combinations of two or three sources. Multiple sources didn't balance out the recommendations, though. Wirecutter tended to dominate for most models whenever it was part of the mix, though the strength of the effect varied. More sources actually led to more variability, according to the study.

Even the order of the same sources changes the outcome
In a third study, agents received all three sources in different orders. A stable decision process should produce the same result given identical content. It didn't.
Gemini 3.1 Flash Lite was the most sensitive, with its probability of choosing the Fitbit Inspire 3 swinging between 2 and 56 percentage points above the control condition depending on source order. Claude Haiku 4.5 stayed stable at 41 to 42 percentage points. The researchers conclude that presentation order is itself a driver of product selection.

Whether sources are passed to the model one at a time or bundled together also matters. GPT-5.5 picked the Fitbit Inspire 3 in 53 percentage points more cases with bundled delivery, but only 6 percentage points more with sequential delivery.
Memory snippets override objective product superiority
In a fourth experiment, the researchers modified the product grid so one product was superior on every measurable dimension: a smart watch with Alexa for $29.99, rated 5.0 out of 5.0 with 430 reviews. Every other product cost at least $359 and had fewer reviews.
Then they added short user memory statements like "I love hiking!" For several models, these statements shifted selections toward pricier products despite the presence of an objectively superior option. The selection rate for the Garmin Vivoactive 5 jumped by 75 percentage points for Claude Opus 4.8, by 37 percentage points for GPT-5.5, and by 36 for Gemini 3.1 Flash Lite.

Gemini 3.5 Flash was the most resistant, picking the objectively best product in 86 to 92 percent of runs regardless of memory statements. GPT-5 Mini showed a strange pattern: the positive hiking statement didn't produce a significant shift toward the Garmin Vivoactive 5, but the negative one ("I don't like hiking!") significantly boosted picks for the Fitbit Versa 4.
No consistency, no control
For shoppers, the study suggests that letting an AI agent buy on your behalf doesn't guarantee consistent or optimal purchase decisions. Two users with the same query, or the same user on a different day, can get different product recommendations with no visible reason. Human buying decisions are inconsistent too, but that's hardly what people expect from an AI shopper. Anyone who has set up a memory in ChatGPT or similar tools should also know that it can affect purchase recommendations in unpredictable ways.
For sellers, the results suggest that optimizing for AI shopping will be harder than traditional SEO, according to the researchers. Sellers don't know which model is doing the shopping, what it read beforehand, or how its technical setup processes information.
Read on for the full picture.
Subscribe for hype-free coverage.
- Full access to every article on THE DECODER
- No ads
- Join the comments and community discussions
- A weekly AI news recap via mail
- 6x/year: "AI Radar" — deep dives on the AI topics that matter most
- Daily AI news, always up to date
- Our full ten-year archive
- Covered by a team with 10+ years in AI