{"id":"15c4f589-c626-4b5d-957c-aec461a23411","arxiv_id":"2508.02634","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Frontier AI shopping agents show strong, model-dependent position biases and unstable market shares, and simple seller-side title edits can shift agent selections by tens of percentage points.","lead":"This paper tests six AI shopping agents in a controlled mock store and finds large, model-specific biases: where a product sits on the page changes how often agents pick it, sponsored tags are penalized, and small seller-side title edits can swing market shares sharply. A generalist reader should care because these behaviors suggest AI-mediated shopping markets are volatile, manipulable, and in need of continuous auditing rather than one-time certification.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No human baseline in the same environment means the 'fundamentally different from human-centric commerce' claim is unsupported; a human-subject replication of the ACES experiments would settle it.","rationale":"The reader identified the sandbox's external validity as the weakest assumption, and my concern overlaps but is more specific: even granting that the sandbox approximates real agentic shopping, the paper's headline claim that agentic markets are 'fundamentally different from human-centric commerce' requires a direct human comparison in the same environment. The reader's condition 4 (add a human baseline or calibrate to an idealized rational chooser) already captures this, so my concern does not change the verdict. I partially disagree with the reader's framing because the reader's weakest_assumption focuses on whether the one-shot, generic-prompt mock store approximates real shopping journeys, whereas my emphasis is on the missing human control as the reference point for 'different.' The seller-agent information advantage is a secondary, related concern: the seller experiments give the optimizer clean counterfactual sales data from the exact buying model, which is not available to real sellers and may inflate the measured gains. Both concerns are addressable with the proposed human-panel replication, which would also provide the missing calibration point for interpreting effect sizes. The paper is otherwise well-executed: the randomizations are clean, the conditional logit specification is standard, and the robustness checks (headless, prompt variations, model drift) materially strengthen the descriptive findings. The conditions the reader imposes are appropriate, and the verdict of CONDITIONAL remains correct.","tokens_in":36976,"tokens_out":10143,"duration_ms":120301,"concrete_test":"Recruit a human panel (e.g., N=200 per category on Prolific), present the same mock-app grid with the same eight products, the same randomizations of position, price, rating, reviews, and badges (Table 1), and ask participants to choose one product under the same generic shopping instruction. Estimate the conditional logit (5.1) on the human choices and compute market-share concentration (e.g., Herfindahl index). If the human position coefficients, sponsored-tag penalty, and Overall Pick lift are statistically indistinguishable from the AI models, the 'fundamentally different' claim fails; if humans show significantly smaller position effects and more dispersed shares, the AI-specific volatility claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 4 treat ACES selection frequencies as market shares and conclude that 'agentic markets are volatile and fundamentally different from human-centric commerce.' Section 8 explicitly defers a human baseline, calling it 'an important question' but outside the paper's aim. This deferral is load-bearing because the central policy and novelty claims depend on AI agents behaving differently from humans. Existing human-subject research documents strong position effects in online rankings (Ursu 2018) and positive responses to platform badges (Lill et al. 2024), which are the same directions as the AI patterns reported here. Without running the identical eight-product grid, generic prompt, and randomized design with human participants, one cannot distinguish an AI-specific distortion from a general property of any decision-maker facing a grid. The seller-side experiments (Section 6.2) compound this: the seller agent is fed 200 clean experimental trials of the exact buying model's choices, an information advantage absent in real markets, so the reported 'significant market share gains' may be upper bounds that do not reflect real seller capabilities. The paper's own robustness checks (prompt variations, headless interface) show that effect sizes are sensitive to the decision context, which further undercuts the unqualified market-share language.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript (arXiv:2508.02630v3, as appears in the full text) introduces ACES, a controllable agent-platform sandbox for auditing AI shopping agents. Using randomized experiments with frontier vision-language models (Claude, GPT, Gemini) in an eight-product grid, the paper measures instruction following, market shares, and causal responses to position, badges, price, ratings, and reviews via conditional logit models. It reports strong model-dependent position biases that persist in headless/API settings, a causal penalty for Sponsored tags, a large positive effect of Overall Pick endorsements, and significant seller-side market-share gains from AI-generated title edits. The paper argues that agentic markets are volatile, model-dependent, and fundamentally different from human-centric commerce, and proposes continuous auditing frameworks. The core empirical design is sound: positions, badges, and attributes are randomized, and the seller experiments reuse identical product shuffles pre/post, enabling causal attribution. However, several claims are overstated relative to the evidence, and the submitted abstract does not match the manuscript's content.","tokens_in":37103,"tokens_out":5822,"duration_ms":63807,"significance":"If the findings hold, this paper is significant for platform design, seller strategy, and AI governance. It provides one of the first rigorous, provider-agnostic frameworks for causally auditing AI shopping agents, with randomized identification of position and badge effects. The headless-interface and prompt-variation robustness checks strengthen the case that these biases are not artifacts of screenshot parsing. The paper ships code and data (GitHub, HuggingFace), supporting replication. The finding that simple title edits can swing market shares—with a concrete mechanism (keyword front-loading)—is a practical and falsifiable contribution. However, the lack of a human baseline and the strong cross-model comparisons limit the external-validity claims.","major_comments":[{"comment":"The abstract at the beginning of the manuscript describes a different paper—actionable counterfactual explanations using Bayesian networks and an EPA dataset—and does not correspond to the full text, which is about AI shopping agents in ACES. This mismatch must be corrected; it is not a minor typo and would mislead readers. The editor should verify that the correct abstract is attached.","section":"Abstract (submitted text)"},{"comment":"The claim that agentic markets are 'fundamentally different from human-centric commerce' is not supported by a direct human baseline. The paper defers a human-subject study (Section 8) but states the stronger claim in the abstract. Existing human literature (e.g., Ursu 2018, Lill et al. 2024) shows similar position and badge effects, so without running the same grid and prompt with human participants, the 'fundamentally different' assertion is an interpretation rather than a measured result. Please soften the claim or add a human baseline.","section":"Abstract and Section 4"},{"comment":"Market-share figures (Figure 1 and Figures 5–6) report point estimates from 200 trials per category without standard errors or confidence intervals. Claims such as 'Fitbit Inspire jumped from 45% to 77%' need uncertainty quantification; with n=200, the standard error of a proportion is roughly 3.5 percentage points, so formal cross-model hypothesis tests should be reported to support the 'drastic swings' language.","section":"Section 4"},{"comment":"The seller agent is given the exact trial-level sales data of the buying agent and full competitor listings, an information advantage that real sellers would not have. The paper should state explicitly that the reported market-share gains are upper bounds under an informed seller, and discuss how the effects might attenuate with realistic, noisier information. This caveat is central to the 'AI-SEO can drive significant gains' contribution.","section":"Section 6.2"}],"minor_comments":[{"comment":"The footnote states that products never selected were excluded from the conditional logit dataset, reducing the average alternatives per choice set to 6.8. This is standard, but the paper should clarify that this does not affect the consistency of the estimates and should report how many products were dropped per model.","section":"Section 5.2, Table 2 note"},{"comment":"In Section 1, the text cites '[13]' for 'race and gender disparities in AI-based hiring', but reference [13] is Gaarlandt et al., 'AI agents are changing how people shop'. This citation appears to point to the wrong source; please verify and correct.","section":"References"},{"comment":"The temporal drift analysis re-ran experiments in September 2025 with the same model names (e.g., Claude Sonnet 4). The paper should clarify whether these are the exact same model checkpoints or whether API updates could have occurred, since model finalization is a studied phenomenon in the paper itself.","section":"Section 7.1"},{"comment":"The conditional logit specification (5.1) assumes IIA; the paper does not discuss this assumption or test its plausibility in the product-choice context. A brief discussion or a robustness check (e.g., nested logit or mixed logit) would strengthen the interpretation of the coefficients.","section":"Section 5.2"},{"comment":"The price-equivalent trade-off for Gemini 3.0 Pro Preview reports a Row 1 premium of +159.7% and an Overall Pick premium of +159.3%. These values are much larger than for other models; please verify the arithmetic and the standard errors.","section":"Table 11"}],"recommendation":"major_revision","confidential_remarks":"The abstract at the top of the submission is for a different paper (Bayesian-network counterfactual explanations) and does not match the full text. Please check whether the wrong manuscript version was uploaded. Also, the arXiv number in the header (2508.02630) differs from the assigned number in the review request (2508.02634); this should be reconciled. The main empirical work is careful and the randomized design is a strength, but the authors should either add a human baseline or retract the 'fundamentally different' claim in the abstract, and they should quantify uncertainty in the market-share comparisons."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things up front. First, the metadata says this is a paper on actionable counterfactual explanations with Bayesian networks, but the actual manuscript is an empirical study of AI shopping agents called 'What Is Your AI Agent Buying?' That identity mismatch is an administrative problem the authors need to resolve before anything else. Second, once you read the body as what it is, the empirical core is genuinely good. What is new: the ACES sandbox and the randomized trials behind it. The design is clean. Positions, badges, prices, ratings, and reviews are all exogenously randomized, and the seller experiments reuse identical shuffles before and after the description edit, so the reported effects are causally attributable. The findings are worth taking seriously: position biases are large, their direction flips between model generations (GPT-4.1 vs GPT-5.1 is a striking example), sponsored tags are penalized, Overall Pick tags produce large lifts, and simple title edits can swing market shares dramatically. The robustness work, including headless API-style interfaces, prompt variations, and the Gemini preview-to-final natural experiment, strengthens the core claims. The authors also released code and data, which is the right move. The soft spots are real but not fatal. The biggest is the absence of a human baseline. The abstract says agentic markets are 'fundamentally different from human-centric commerce,' but the paper never runs the same grid with human participants. Existing human-subject work shows strong position effects and positive badge responses, so the AI behaviors may be less novel than claimed. Section 8 explicitly defers this, which is honest, but the abstract and conclusions have to be tempered. Second, the conditional logit drops never-selected products, which is standard practice but worth disclosing more prominently because it can bias position and attribute coefficients. Third, the seller-side experiments give the seller agent 200 clean trials of the exact buying model's choices, a clear information advantage over real markets, so the market-share gains should be framed as upper bounds. Fourth, the 'improvement over time' language is really a cross-generation comparison, not a longitudinal trend, and model snapshots may not generalize. If this body is what the authors intend to submit, it deserves serious refereeing after the framing is fixed. As a reviewer, I would ask for the human baseline or a downgraded claim, a fuller treatment of the selection issue, and a resolution of the metadata mismatch. As declared on arXiv, the identity problem alone would justify rejection until corrected. My vote: engage with the body, not the metadata.","headline":"The body under review is a solid, carefully randomized measurement study of AI shopping agents, but it is not the paper the arXiv metadata advertises, and its headline claims outrun the evidence by skipping a human baseline.","tokens_in":750,"tokens_out":934,"would_cite":true,"duration_ms":28915,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Frontier AI shopping agents exhibit large position biases that persist across interfaces and invert when models are upgraded.","keywords":["AI shopping agents","agentic e-commerce","position bias","sponsored tags","platform endorsements","market share volatility","conditional logit","seller-side optimization"],"falsifier":"Run the same position-randomization trials with an agent that browses multiple pages, opens product-detail pages, and carries user context; if the position bias and sponsored-tag penalty shrink to near zero under full-funnel browsing, the claim that these distortions are structural features of agentic commerce would fail. A second test: a released model generation that preserves its predecessor's position preferences would falsify the claim that upgrades systematically invert or reshuffle these biases.","tokens_in":36618,"feed_emoji":"🛒","tokens_out":6825,"duration_ms":67427,"temperature":0.7,"pith_summary":"This paper tries to establish that autonomous AI shopping agents are a new kind of economic decision-maker with systematic, exploitable biases, rather than the rational optimizers economic theory often assumes. Using randomized experiments in a controllable mock store, it shows that where a product sits on a page can change its selection probability several-fold, that the preferred slot can flip between model generations, and that a 'Sponsored' tag depresses choice while an 'Overall Pick' endorsement strongly lifts it. It also shows that sellers can exploit these patterns: simple, query-conditional title edits produced statistically significant market-share gains for the focal product in five of six buyer models tested. If true, AI-mediated commerce is volatile, manipulable, and different in kind from human-centric shopping, which matters for platform design, seller strategy, product rankings, advertising, and regulation.","feed_headline":"AI shoppers favor slots, not products — and upgrades flip the winners","feed_subtitle":"Randomized trials show sponsored tags hurt agent choices, endorsements help, and simple title edits swing market share.","key_machinery":"ACES (Agentic e-Commerce Simulator): a provider-agnostic evaluation sandbox that pairs a vision-language agent, controlling a browser through tools, with a programmable mock storefront rendering an eight-product grid in a randomized layout. Its engine is the randomized controlled trial: positions, tag assignments, prices, ratings, and review counts are exogenously varied, and a conditional logit model recovers causal choice probabilities and price-equivalent trade-offs between levers. The same trials are reproduced in a headless variant that exposes only a ranked JSON list with no images, which separates visual parsing artifacts from deep-seated model priors.","core_discovery":"The paper's central claim is that frontier AI shopping agents choose products in ways that are simultaneously homogeneous, unstable, and causally responsive to platform levers. Across hundreds of randomized trials per product category, selection shares collapse onto a few modal products while other brands are never chosen; model upgrades reshuffle shares drastically, with the Fitbit Inspire's share jumping from 25% to 77% after one Claude upgrade and falling from 25% to 6% after a GPT upgrade. Conditional-logit estimates show position effects large enough that moving a product from the bottom-right slot to the top row can raise its selection rate five-fold, and the preferred position of GPT-4.1 is the least preferred position of its successor GPT-5.1. Because badge assignment and position were randomized, the paper reads the sponsored-tag penalty and the Overall Pick lift as causal, with the endorsement worth as much as a 65-138% price increase in price-equivalent terms. A seller-side agent making one-shot, query-conditional description edits produced significant share gains in five of six buyer models, with office-lamp gains of up to 80.4 percentage points from front-loading the word 'Office'.","pith_inferences":["A testable extension is to run the same position-randomization trials with agents that browse multiple pages, open product-detail pages, and carry user context; if position effects attenuate under full-funnel browsing, the 'deep-seated characteristic' reading would need to be revised downward.","The price-equivalent trade-offs the paper computes suggest that platforms could price endorsements, premium slots, and even listing-text services in an 'agent currency', a design implication the paper states only implicitly.","The inverted position preferences between adjacent model generations hint that post-training choices, rather than the pretrained knowledge base, drive much of the bias; direct attribution experiments on model checkpoints could test this.","If choice homogeneity persists as agents scale, marketplaces may need diversity-promoting mechanisms or neutrality audits, since AI-mediated demand could otherwise produce winner-take-all outcomes regardless of underlying consumer taste."],"forward_implications":["If AI agents mediate a growing share of purchases, product rankings become a first-order market lever, and there is no universal 'top' slot because the best position is model-specific.","Platform endorsements like Overall Pick act as powerful credibility signals to agents while sponsored tags carry a credibility penalty, shifting how advertising and promotion budgets should be valued.","Model updates function as exogenous demand shocks, so sellers with static listings can gain or lose market share overnight even when catalogs are unchanged.","Simple seller-side text optimization can unlock large share gains, implying that listing titles will become a strategic battlefield and that static descriptions are suboptimal against algorithmic buyers.","Because position effects persist in headless/API settings and resist 'ignore position' prompting, the distortions are unlikely to disappear without deliberate design intervention.","The paper's finding that agents concentrate choice on a few modal products implies concentration risk: dominant agents could suppress niche brands that would compete normally under human demand."],"supporting_citations":[{"why":"Supplies the conditional logit model used to estimate causal choice sensitivities to positions, tags, and attributes from the randomized trials.","marker":"[32]"},{"why":"Provides the empirical baseline showing rankings causally shape human search and purchase decisions, the human counterpart to the agent position effects.","marker":"[41]"},{"why":"Documents causal effects of product badges on human choice, the comparison point for the agent responses to Overall Pick, Sponsored, and scarcity tags.","marker":"[29]"},{"why":"Quantifies how rankings affect consumer behavior and platform revenue, framing the platform-design stakes the agent results extend.","marker":"[14]"},{"why":"Supplies human price-elasticity benchmarks used to interpret the magnitude of the agent price coefficients.","marker":"[4]"},{"why":"The computer-use testbed whose end-to-end task evaluation ACES deliberately narrows to isolate the single product-selection step.","marker":"[50]"},{"why":"The multimodal web-agent benchmark that motivates the vision-language screenshot interface used in the sandbox.","marker":"[26]"},{"why":"Documents rapid adoption of AI shopping assistants, the market condition that motivates why agent choice behavior matters now.","marker":"[13]"}],"fun_headline_variants":["AI shoppers fixate on slots, not brands","Model upgrades flip AI shopping selections","Sponsored tags hurt AI agents, endorsements help","Position effects five-fold AI shopping pick rates","AI agents cluster on few products, ignoring rest"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a one-shot choice by an unassisted vision-language agent over an eight-product grid, from a generic prompt in a synthetic mock store, faithfully approximates real agentic shopping, with sampled selection frequencies treated as market shares.","fun_headline_variants_meta":{"raw":{"variants":["AI shoppers fixate on slots, not brands","Model upgrades flip AI shopping selections","Sponsored tags hurt AI agents, endorsements help","Position effects five-fold AI shopping pick rates","AI agents cluster on few products, ignoring rest"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000328,"raw_usage":{"total_tokens":1881,"prompt_tokens":1041,"completion_tokens":840,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":657,"completion_tokens_details":{"reasoning_tokens":773}},"tokens_in":657,"tokens_out":840,"duration_ms":8047,"temperature":1.0,"reasoning_tokens":773,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:39:59.682018+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same position-randomization trials with an agent that browses multiple pages, opens product-detail pages, and carries user context; if the position bias and sponsored-tag penalty shrink to near zero under full-funnel browsing, the claim that these distortions are structural features of agentic commerce would fail. A second test: a released model generation that preserves its predecessor's position preferences would falsify the claim that upgrades systematically invert or reshuffle these biases.","supporting_citations":[{"cited_title":"Conditional logit analysis of qualitative choice behavior","cited_arxiv_id":null,"evidence_quote":"Supplies the conditional logit model used to estimate causal choice sensitivities to positions, tags, and attributes from the randomized trials."},{"cited_title":"The power of rankings: Quantifying the effect of rankings on online con- sumer search and purchase decisions.��������� �������, 37(4):530–552, 2018","cited_arxiv_id":null,"evidence_quote":"Provides the empirical baseline showing rankings causally shape human search and purchase decisions, the human counterpart to the agent position effects."},{"cited_title":"Product badges and consumer choice on digital platforms.��������� �� ���� �������, 2024","cited_arxiv_id":null,"evidence_quote":"Documents causal effects of product badges on human choice, the comparison point for the agent responses to Overall Pick, Sponsored, and scarcity tags."},{"cited_title":"Examining the impact of ranking on consumer behavior and search engine revenue.���������� �������, 60(7):1632–1654, 2014","cited_arxiv_id":null,"evidence_quote":"Quantifies how rankings affect consumer behavior and platform revenue, framing the platform-design stakes the agent results extend."},{"cited_title":"Bijmolt, Harald J","cited_arxiv_id":null,"evidence_quote":"Supplies human price-elasticity benchmarks used to interpret the magnitude of the agent price coefficients."},{"cited_title":"Webarena: A realistic web environment for building autonomous agents.����� �������� ����������������, 2024","cited_arxiv_id":null,"evidence_quote":"The computer-use testbed whose end-to-end task evaluation ACES deliberately narrows to isolate the single product-selection step."},{"cited_title":"Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.����� �������� ����������������,","cited_arxiv_id":null,"evidence_quote":"The multimodal web-agent benchmark that motivates the vision-language screenshot interface used in the sandbox."},{"cited_title":"Ai agents are changing how people shop","cited_arxiv_id":null,"evidence_quote":"Documents rapid adoption of AI shopping assistants, the market condition that motivates why agent choice behavior matters now."}],"review_version":2}