메뉴
BL
The Decoder • 29일 전

연구 결과, AI 쇼핑 에이전트는 아직 대리 구매를 맡기기엔 이르다

IMP
7/10
핵심 요약

펜실베이니아대 와튼스쿨 연구진이 6개 최신 AI 모델을 쇼핑 에이전트로 테스트한 결과, 외부 출처 하나만 추가되거나 제시 순서만 바뀌어도 제품 추천이 크게 달라지는 것으로 나타났다. 심지어 객관적으로 우월한 제품이 있어도 사용자 기억 메모 한 줄에 추천이 고가 제품으로 쏠리는 등 일관성이 결여됐다. 사용자에게는 보이지 않는 이유로 같은 검색이라도 매번 다른 추천이 나올 수 있어, AI에 구매를 위임하기엔 아직 이르다는 결론이다.

번역된 본문

연구 결과, AI 쇼핑 에이전트는 아직 사용자를 대신해 구매할 준비가 되지 않았다

Matthias Bastian, 2026년 8월 27일, THE DECODER

펜실베이니아대 와튼스쿨의 연구진은 검색 과정이 달라질 때 AI 쇼핑 에이전트가 제품을 얼마나 일관되게 추천하는지 테스트했다. 그 결과, 아주 사소한 맥락 변화만으로도 구매 결정이 크게 흔들리는 것으로 나타났다.

연구팀은 미니 버전부터 최상위(frontier) 모델까지 현재 사용 가능한 6개 모델을 테스트했으며, 각 모델에게 고정된 제품 목록에서 피트니스 워치를 고르는 개인 쇼핑 비서 역할을 맡겼다. 이들은 AI 에이전트에게 제품 페이지 스크린샷을 보여주는 ACES 시뮬레이터(Agentic e-Commerce Simulator)를 활용했다. 에이전트는 이미지를 분석하고, 선택적으로 추천 출처를 참고한 뒤 제품을 하나 고른다.

하나의 출처만으로도 추천이 뒤집힌다

외부 출처 없이도 모델들은 서로 다른 기본 선호를 보였다. 하지만 에이전트가 제품 페이지를 보기 전 단 하나의 외부 출처만 접해도 경우에 따라 추천이 극적으로 바뀌었다. 연구진은 세 가지 출처를 테스트했다: 가민 포러너 55를 추천하는 레딧 스레드, 핏빗 인스파이어 3에 대한 와이어커터 리뷰, WHOOP 5.0에 관한 스트래티스트 기사다. 와이어커터의 영향력이 가장 컸다. 핏빗 인스파이어 3를 선택할 확률은 대조군 대비 Claude Opus 4.8에서 90%포인트, Gemini 3.5 Flash에서 99%포인트 급등했다.

두 번째 실험에서는 에이전트가 두세 개 출처의 조합을 보게 했다. 하지만 여러 출처가 추천을 균형 있게 만들지는 못했다. 와이어커터가 조합에 포함되기만 하면 대부분의 모델에서 와이어커터가 지배하는 경향을 보였으며, 그 영향력의 크기는 모델마다 달랐다. 연구에 따르면 출처가 많을수록 오히려 변동성도 커졌다.

같은 출처라도 순서만 바꾸면 결과가 달라진다

세 번째 실험에서 에이전트는 세 출처를 서로 다른 순서로 전달받았다. 안정적인 의사결정 과정이라면 동일한 내용에 대해서는 같은 결과가 나와야 한다. 하지만 그렇지 않았다. Gemini 3.1 Flash Lite가 가장 민감해서, 출처 순서에 따라 핏빗 인스파이어 3를 선택할 확률이 대조군 대비 2%포인트에서 56%포인트까지 요동쳤다. Claude Haiku 4.5는 41~42%포인트로 안정적이었다. 연구진은 출처의 제시 순서 자체가 제품 선택을 좌우하는 요인이라고 결론지었다. 출처를 한 번에 하나씩 전달하느냐 묶어서 전달하느냐도 중요했다. GPT-5.5는 묶어서 전달할 때는 53%포인트 더 많이 핏빗 인스파이어 3를 골랐지만, 순차적으로 전달할 때는 6%포인트에 그쳤다.

기억 메모가 객관적 제품 우위를 뒤엎는다

네 번째 실험에서 연구진은 제품 목록을 수정해 한 제품이 모든 측정 가능한 지표에서 우월하도록 만들었다. 29.99달러의 알렉사 내장 스마트워치로, 5점 만점에 5.0점, 리뷰 430개를 보유했다. 나머지 제품은 모두 최소 359달러에 리뷰 수도 더 적었다. 그런 다음 "나는 하이킹을 좋아해!" 같은 짧은 사용자 기억 문장을 추가했다. 여러 모델에서 이런 문장이 객관적으로 우월한 선택지가 있음에도 불구하고 추천을 더 비싼 제품 쪽으로 옮겼다. 가민 비보액티브 5의 선택률은 Claude Opus 4.8에서 75%포인트, GPT-5.5에서 37%포인트, Gemini 3.1 Flash Lite에서 36%포인트 급등했다. Gemini 3.5 Flash가 가장 저항력이 강해, 기억 문장과 무관하게 실행의 86~92%에서 객관적으로 최고인 제품을 골랐다.

GPT-5 Mini는 기묘한 패턴을 보였다. 긍정적인 하이킹 문장은 가민 비보액티브 5로의 유의미한 이동을 일으키지 않았지만, 부정적 문장("나는 하이킹을 좋아하지 않아!")은 핏빗 베르사 4 선택을 유의미하게 끌어올렸다.

일관성도, 통제도 없다

소비자 입장에서 이 연구는 AI 에이전트에게 대리 구매를 맡긴다고 해서 일관되거나 최적의 구매 결정이 보장되지 않는다는 것을 시사한다. 같은 검색어를 쓴 두 사용자, 혹은 같은 사용자가 다른 날 검색해도 눈에 보이는 이유 없이 다른 제품 추천을 받을 수 있다. 인간의 구매 결정 역시...

원문 보기
원문 보기 (영어)
AI shopping agents aren't ready to buy on your behalf, study finds Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 27, 2026 GPT-Image-2 prompted by THE DECODER Researchers at the Wharton School at the University of Pennsylvania tested how consistently AI shopping agents recommend products when the search process changes. Even tiny shifts in context swung purchase decisions by wide margins. The team tested six current models, both mini variants and frontier-level, tasking each one to act as a personal shopping assistant picking a fitness watch from a fixed product grid. They used the ACES simulator (Agentic e-Commerce Simulator), which shows the AI agent a screenshot of a product page. The agent analyzes the image, optionally pulls in recommendation sources, and then picks a product. A single source is enough to flip the recommendation Even without external sources, the models showed different baseline preferences. But when the agent saw just one external source before the product page, recommendations shifted dramatically in some cases . The researchers tested three sources: a Reddit thread recommending the Garmin Forerunner 55, a Wirecutter review for the Fitbit Inspire 3, and a Strategist article about the WHOOP 5.0. Wirecutter had the strongest pull. The probability of picking the Fitbit Inspire 3 jumped by 90 percentage points for Claude Opus 4.8 compared to the control condition, and by 99 percentage points for Gemini 3.5 Flash. In a second experiment, agents saw combinations of two or three sources. Multiple sources didn't balance out the recommendations, though. Wirecutter tended to dominate for most models whenever it was part of the mix, though the strength of the effect varied. More sources actually led to more variability, according to the study. Even the order of the same sources changes the outcome In a third study, agents received all three sources in different orders. A stable decision process should produce the same result given identical content. It didn't. Gemini 3.1 Flash Lite was the most sensitive, with its probability of choosing the Fitbit Inspire 3 swinging between 2 and 56 percentage points above the control condition depending on source order. Claude Haiku 4.5 stayed stable at 41 to 42 percentage points. The researchers conclude that presentation order is itself a driver of product selection. Whether sources are passed to the model one at a time or bundled together also matters. GPT-5.5 picked the Fitbit Inspire 3 in 53 percentage points more cases with bundled delivery, but only 6 percentage points more with sequential delivery. Memory snippets override objective product superiority In a fourth experiment, the researchers modified the product grid so one product was superior on every measurable dimension: a smart watch with Alexa for $29.99, rated 5.0 out of 5.0 with 430 reviews. Every other product cost at least $359 and had fewer reviews. Then they added short user memory statements like "I love hiking!" For several models, these statements shifted selections toward pricier products despite the presence of an objectively superior option. The selection rate for the Garmin Vivoactive 5 jumped by 75 percentage points for Claude Opus 4.8, by 37 percentage points for GPT-5.5, and by 36 for Gemini 3.1 Flash Lite. Gemini 3.5 Flash was the most resistant, picking the objectively best product in 86 to 92 percent of runs regardless of memory statements. GPT-5 Mini showed a strange pattern: the positive hiking statement didn't produce a significant shift toward the Garmin Vivoactive 5, but the negative one ("I don't like hiking!") significantly boosted picks for the Fitbit Versa 4. No consistency, no control For shoppers, the study suggests that letting an AI agent buy on your behalf doesn't guarantee consistent or optimal purchase decisions. Two users with the same query, or the same user on a different day, can get different product recommendations with no visible reason. Human buying decisions are inconsistent too, but that's hardly what people expect from an AI shopper. Anyone who has set up a memory in ChatGPT or similar tools should also know that it can affect purchase recommendations in unpredictable ways. For sellers, the results suggest that optimizing for AI shopping will be harder than traditional SEO, according to the researchers. Sellers don't know which model is doing the shopping, what it read beforehand, or how its technical setup processes information. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->