Pith. sign in

REVIEW 4 major objections 5 minor 61 references

Actionable Counterfactual Explanations Using Bayesian Networks and Path Planning with Applications to Environmental Quality Improvement

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Frontier AI shopping agents exhibit large position biases that persist across interfaces and invert when models are upgraded.

desk verdict The body under review is a solid, carefully randomized measurement study of AI shopping agents, but it is not the paper the arXiv metadata advertises, and its headline claims outrun the evidence by skipping a human baseline. read the letter →

arxiv 2508.02634 v1 pith:4JIKVWK5 submitted 2025-08-04 cs.AI cs.LG

classification cs.AIcs.LG
keywords AIshoppingagentsagentice-commercepositionbiassponsoredtagsplatformendorsementsmarketsharevolatilityconditionallogitseller-sideoptimization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that autonomous AI shopping agents are a new kind of economic decision-maker with systematic, exploitable biases, rather than the rational optimizers economic theory often assumes. Using randomized experiments in a controllable mock store, it shows that where a product sits on a page can change its selection probability several-fold, that the preferred slot can flip between model generations, and that a 'Sponsored' tag depresses choice while an 'Overall Pick' endorsement strongly lifts it. It also shows that sellers can exploit these patterns: simple, query-conditional title edits produced statistically significant market-share gains for the focal product in five of six buyer models tested. If true, AI-mediated commerce is volatile, manipulable, and different in kind from human-centric shopping, which matters for platform design, seller strategy, product rankings, advertising, and regulation.

What carries the argument

ACES (Agentic e-Commerce Simulator): a provider-agnostic evaluation sandbox that pairs a vision-language agent, controlling a browser through tools, with a programmable mock storefront rendering an eight-product grid in a randomized layout. Its engine is the randomized controlled trial: positions, tag assignments, prices, ratings, and review counts are exogenously varied, and a conditional logit model recovers causal choice probabilities and price-equivalent trade-offs between levers. The same trials are reproduced in a headless variant that exposes only a ranked JSON list with no images, which separates visual parsing artifacts from deep-seated model priors.

What would settle it

Run the same position-randomization trials with an agent that browses multiple pages, opens product-detail pages, and carries user context; if the position bias and sponsored-tag penalty shrink to near zero under full-funnel browsing, the claim that these distortions are structural features of agentic commerce would fail. A second test: a released model generation that preserves its predecessor's position preferences would falsify the claim that upgrades systematically invert or reshuffle these biases.

Watch

Extended reading notes

Core claim

The paper's central claim is that frontier AI shopping agents choose products in ways that are simultaneously homogeneous, unstable, and causally responsive to platform levers. Across hundreds of randomized trials per product category, selection shares collapse onto a few modal products while other brands are never chosen; model upgrades reshuffle shares drastically, with the Fitbit Inspire's share jumping from 25% to 77% after one Claude upgrade and falling from 25% to 6% after a GPT upgrade. Conditional-logit estimates show position effects large enough that moving a product from the bottom-right slot to the top row can raise its selection rate five-fold, and the preferred position of GPT-4.1 is the least preferred position of its successor GPT-5.1. Because badge assignment and position were randomized, the paper reads the sponsored-tag penalty and the Overall Pick lift as causal, with the endorsement worth as much as a 65-138% price increase in price-equivalent terms. A seller-side agent making one-shot, query-conditional description edits produced significant share gains in five of six buyer models, with office-lamp gains of up to 80.4 percentage points from front-loading the word 'Office'.

Load-bearing premise

The load-bearing premise is that a one-shot choice by an unassisted vision-language agent over an eight-product grid, from a generic prompt in a synthetic mock store, faithfully approximates real agentic shopping, with sampled selection frequencies treated as market shares.

Editorial extensions

If this is right

  • If AI agents mediate a growing share of purchases, product rankings become a first-order market lever, and there is no universal 'top' slot because the best position is model-specific.
  • Platform endorsements like Overall Pick act as powerful credibility signals to agents while sponsored tags carry a credibility penalty, shifting how advertising and promotion budgets should be valued.
  • Model updates function as exogenous demand shocks, so sellers with static listings can gain or lose market share overnight even when catalogs are unchanged.
  • Simple seller-side text optimization can unlock large share gains, implying that listing titles will become a strategic battlefield and that static descriptions are suboptimal against algorithmic buyers.
  • Because position effects persist in headless/API settings and resist 'ignore position' prompting, the distortions are unlikely to disappear without deliberate design intervention.
  • The paper's finding that agents concentrate choice on a few modal products implies concentration risk: dominant agents could suppress niche brands that would compete normally under human demand.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to run the same position-randomization trials with agents that browse multiple pages, open product-detail pages, and carry user context; if position effects attenuate under full-funnel browsing, the 'deep-seated characteristic' reading would need to be revised downward.
  • The price-equivalent trade-offs the paper computes suggest that platforms could price endorsements, premium slots, and even listing-text services in an 'agent currency', a design implication the paper states only implicitly.
  • The inverted position preferences between adjacent model generations hint that post-training choices, rather than the pretrained knowledge base, drive much of the bias; direct attribution experiments on model checkpoints could test this.
  • If choice homogeneity persists as agents scale, marketplaces may need diversity-promoting mechanisms or neutrality audits, since AI-mediated demand could otherwise produce winner-take-all outcomes regardless of underlying consumer taste.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript (arXiv:2508.02630v3, as appears in the full text) introduces ACES, a controllable agent-platform sandbox for auditing AI shopping agents. Using randomized experiments with frontier vision-language models (Claude, GPT, Gemini) in an eight-product grid, the paper measures instruction following, market shares, and causal responses to position, badges, price, ratings, and reviews via conditional logit models. It reports strong model-dependent position biases that persist in headless/API settings, a causal penalty for Sponsored tags, a large positive effect of Overall Pick endorsements, and significant seller-side market-share gains from AI-generated title edits. The paper argues that agentic markets are volatile, model-dependent, and fundamentally different from human-centric commerce, and proposes continuous auditing frameworks. The core empirical design is sound: positions, badges, and attributes are randomized, and the seller experiments reuse identical product shuffles pre/post, enabling causal attribution. However, several claims are overstated relative to the evidence, and the submitted abstract does not match the manuscript's content.

Significance. If the findings hold, this paper is significant for platform design, seller strategy, and AI governance. It provides one of the first rigorous, provider-agnostic frameworks for causally auditing AI shopping agents, with randomized identification of position and badge effects. The headless-interface and prompt-variation robustness checks strengthen the case that these biases are not artifacts of screenshot parsing. The paper ships code and data (GitHub, HuggingFace), supporting replication. The finding that simple title edits can swing market shares—with a concrete mechanism (keyword front-loading)—is a practical and falsifiable contribution. However, the lack of a human baseline and the strong cross-model comparisons limit the external-validity claims.

major comments (4)
  1. [Abstract (submitted text)] The abstract at the beginning of the manuscript describes a different paper—actionable counterfactual explanations using Bayesian networks and an EPA dataset—and does not correspond to the full text, which is about AI shopping agents in ACES. This mismatch must be corrected; it is not a minor typo and would mislead readers. The editor should verify that the correct abstract is attached.
  2. [Abstract and Section 4] The claim that agentic markets are 'fundamentally different from human-centric commerce' is not supported by a direct human baseline. The paper defers a human-subject study (Section 8) but states the stronger claim in the abstract. Existing human literature (e.g., Ursu 2018, Lill et al. 2024) shows similar position and badge effects, so without running the same grid and prompt with human participants, the 'fundamentally different' assertion is an interpretation rather than a measured result. Please soften the claim or add a human baseline.
  3. [Section 4] Market-share figures (Figure 1 and Figures 5–6) report point estimates from 200 trials per category without standard errors or confidence intervals. Claims such as 'Fitbit Inspire jumped from 45% to 77%' need uncertainty quantification; with n=200, the standard error of a proportion is roughly 3.5 percentage points, so formal cross-model hypothesis tests should be reported to support the 'drastic swings' language.
  4. [Section 6.2] The seller agent is given the exact trial-level sales data of the buying agent and full competitor listings, an information advantage that real sellers would not have. The paper should state explicitly that the reported market-share gains are upper bounds under an informed seller, and discuss how the effects might attenuate with realistic, noisier information. This caveat is central to the 'AI-SEO can drive significant gains' contribution.
minor comments (5)
  1. [Section 5.2, Table 2 note] The footnote states that products never selected were excluded from the conditional logit dataset, reducing the average alternatives per choice set to 6.8. This is standard, but the paper should clarify that this does not affect the consistency of the estimates and should report how many products were dropped per model.
  2. [References] In Section 1, the text cites '[13]' for 'race and gender disparities in AI-based hiring', but reference [13] is Gaarlandt et al., 'AI agents are changing how people shop'. This citation appears to point to the wrong source; please verify and correct.
  3. [Section 7.1] The temporal drift analysis re-ran experiments in September 2025 with the same model names (e.g., Claude Sonnet 4). The paper should clarify whether these are the exact same model checkpoints or whether API updates could have occurred, since model finalization is a studied phenomenon in the paper itself.
  4. [Section 5.2] The conditional logit specification (5.1) assumes IIA; the paper does not discuss this assumption or test its plausibility in the product-choice context. A brief discussion or a robustness check (e.g., nested logit or mixed logit) would strengthen the interpretation of the coefficients.
  5. [Table 11] The price-equivalent trade-off for Gemini 3.0 Pro Preview reports a Row 1 premium of +159.7% and an Overall Pick premium of +159.3%. These values are much larger than for other models; please verify the arithmetic and the standard errors.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: an empirical measurement paper whose headline quantities are direct experimental outcomes, not derived predictions or renamed inputs.

full rationale

The paper does not derive its headline claims from a fitted model, a self-citation, or a uniqueness theorem. Market shares are direct selection frequencies from randomized trials (Section 4.1: 'we compute the induced market shares, representing the selection rates of different products across the 400 experiments'). The conditional logit estimates in Section 5.2 are descriptive summaries of the same experimental choices, and the position-probability plots are explicitly labeled 'model-based estimates' for illustration, not independent predictions. Seller-side gains are measured as observed pre/post differences in selection probability under identical shuffles (Section 6.2: 'we reuse the identical product shuffles in the pre/post modification experiments'), so the outcome is experimental rather than derived from the fitted model. The headless interface and prompt-robustness checks are separate runs with new choices. No load-bearing premise is justified by a self-citation, and no prediction is a renamed experimental input. The absence of a human baseline is an external-validity limitation, not a circularity.

Assumptions & free parameters 6 free parameters · 4 assumptions · 1 invented entities

This is an empirical measurement paper, so the honest ledger consists of openly fitted choice-model coefficients (the measured effects themselves) plus the modeling and domain assumptions that shape what is measured. The fitted coefficients are not hidden: they are the results reported in Tables 2, 3, and 12-19. The axioms record the main structural choices: the conditional logit/IIA specification, the one-shot mock-store approximation of real shopping, the cost/latency-constrained model configurations, and the author-selected product assortments. No speculative physical or conceptual entities are introduced; the only new construct is the ACES sandbox, which ships as code.

free parameters (6)
  • Conditional logit price coefficient = estimated per model; Tables 2-3, 12-19
    Utility weight on price fitted to the randomized choice data; drives the price-sensitivity and price-equivalent trade-off claims.
  • Conditional logit rating coefficient = estimated per model; Tables 2-3, 12-19
    Utility weight on average rating fitted to choice data.
  • Conditional logit number-of-reviews coefficient = estimated per model; Tables 2-3, 12-19
    Utility weight on review count fitted to choice data.
  • Position dummies (Row 1, Columns 1-3, or Positions 1-7 in headless) = estimated per model; Tables 2-3, 17-19
    These coefficients are the measured position bias, the paper's headline result.
  • Badge dummies (Sponsored, Overall Pick, Scarcity) = estimated per model; Tables 2-3
    Causal effects of randomly assigned tags on agent choice.
  • Product fixed effects = estimated per product per model
    Control for intrinsic product attractiveness in the conditional logit.
assumptions (4)
  • domain assumption Conditional logit with linear index and IIA (Eq. 5.1-5.2) captures agent choice behavior
    Standard discrete-choice specification imposed on agent choices; IIA and linearity are not derived from agent behavior.
  • domain assumption One-shot choice over an eight-product grid with a generic prompt approximates real agentic shopping (Section 2)
    The paper brackets full-funnel browsing, personalization, and multi-step search to isolate the choice step.
  • domain assumption Model configuration (temperature 1.0, minimal reasoning budgets) is a valid deployment baseline (Table 5)
    Biases are measured under cost and latency constrained settings; higher reasoning effort could change behavior.
  • domain assumption Author-selected eight-product assortments from Amazon listings are representative (Section 2)
    Results such as the SUNMORY title-truncation effect depend on these specific listings and their rendering.
invented entities (1)
  • ACES (Agentic e-Commerce Simulator) independent evidence
    purpose: Provider-agnostic sandbox for causally auditing AI shopping-agent choices
    Shipped as source code on GitHub with datasets claimed on HuggingFace, so it is independently usable; it is a measurement tool, not a speculative entity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Actionable Counterfactual Explanations Using Bayesian Networks and Path Planning with Applications to Environmental Quality Improvement." pith.science (2026). https://pith.science/paper/4JIKVWK5

@misc{pith2026250802634,
  author       = {Pith},
  title        = {Pith review of: Actionable Counterfactual Explanations Using Bayesian Networks and Path Planning with Applications to Environmental Quality Improvement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4JIKVWK5}},
  note         = {Machine review of arXiv:2508.02634}
}
read the original abstract

Counterfactual explanations study what should have changed in order to get an alternative result, enabling end-users to understand machine learning mechanisms with counterexamples. Actionability is defined as the ability to transform the original case to be explained into a counterfactual one. We develop a method for actionable counterfactual explanations that, unlike predecessors, does not directly leverage training data. Rather, data is only used to learn a density estimator, creating a search landscape in which to apply path planning algorithms to solve the problem and masking the endogenous data, which can be sensitive or private. We put special focus on estimating the data density using Bayesian networks, demonstrating how their enhanced interpretability is useful in high-stakes scenarios in which fairness is raising concern. Using a synthetic benchmark comprised of 15 datasets, our proposal finds more actionable and simpler counterfactuals than the current state-of-the-art algorithms. We also test our algorithm with a real-world Environmental Protection Agency dataset, facilitating a more efficient and equitable study of policies to improve the quality of life in United States of America counties. Our proposal captures the interaction of variables, ensuring equity in decisions, as policies to improve certain domains of study (air, water quality, etc.) can be detrimental in others. In particular, the sociodemographic domain is often involved, where we find important variables related to the ongoing housing crisis that can potentially have a severe negative impact on communities.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

61 extracted references · 61 canonical work pages

  1. [1]

    Agent s2: A compositional generalist-specialist framework for computer use agents.����� �������� ����������������, 2025

    Saaket Agashe, Kyle Wong, Vincent Tu, Jiachen Yang, Ang Li, and Xin Eric Wang. Agent s2: A compositional generalist-specialist framework for computer use agents.����� �������� ����������������, 2025. (Cited on pages 8 and 37)

  2. [2]

    A model of delegated project choice.������������, 78(1):213–244, 2010

    Mark Armstrong and John Vickers. A model of delegated project choice.������������, 78(1):213–244, 2010. (Cited on page 9)

  3. [3]

    These deals won’t last! longevity, uniformity and bias in product badge assignment in e-commerce platforms.����� �������� ����������������, 2022

    Archit Bansal, Kunal Banerjee, and Abhijnan Chakraborty. These deals won’t last! longevity, uniformity and bias in product badge assignment in e-commerce platforms.����� �������� ����������������, 2022. (Cited on page 8)

  4. [4]

    Bijmolt, Harald J

    Tammo H.A. Bijmolt, Harald J. Van Heerde, and Rik G.M. Pieters. New empirical generaliza- tions on the determinants of price elasticity.������� �� ��������� ��������, 42(2):141–156, 2005. (Cited on page 18) 31

  5. [5]

    Win- dows agent arena: Evaluating multi-modal os agents at scale.����� �������� ����������������,

    Rogerio Bonatti, David Zhao, Francesco Bonacci, David Dupont, Sasan Abdali, Yifan Li, Yix- uan Lu, Jignesh Wagle, Kazuhito Koishida, Andres Bucker, Leon Jang, and Zihang Hui. Win- dows agent arena: Evaluating multi-modal os agents at scale.����� �������� ����������������,

  6. [6]

    Using llms for market research.������� �������� ������ ��������� ���� ������� �����, (23-062), 2023

    James Brand, Ayelet Israeli, and Donald Ngwe. Using llms for market research.������� �������� ������ ��������� ���� ������� �����, (23-062), 2023. (Cited on page 9)

  7. [7]

    Online search and product rankings: A double logit approach

    Giovanni Compiani, Gregory Lewis, Sida Peng, and Peichun Wang. Online search and product rankings: A double logit approach. Technical report, Working paper, University of Chicago, Chicago, IL, 2022. (Cited on page 8)

  8. [8]

    A shopping agent for addressing subjective product needs

    Preetam Prabhu Srikar Dammu, Omar Alonso, and Barbara Poblete. A shopping agent for addressing subjective product needs. In����������� �� ��� ���������� ��� ������������� ���� ������� �� ��� ������ ��� ���� ������, pages 1032–1035, 2025. (Cited on page 8)

Show all 61 references
  1. [9]

    Assortment optimization under variants of the nested logit model.���������� ��������, 62(2):250–273, 2014

    James M Davis, Guillermo Gallego, and Huseyin Topaloglu. Assortment optimization under variants of the nested logit model.���������� ��������, 62(2):250–273, 2014. (Cited on page 8)

  2. [10]

    Mind2web: Towards a generalist agent for the web.�������� �� ������ ����������� ���������� �������, 36:28091–28114, 2023

    Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web.�������� �� ������ ����������� ���������� �������, 36:28091–28114, 2023. (Cited on pages 7 and 31)

  3. [11]

    Product rank- ing on online platforms.���������� �������, 68(6):4024–4041, 2022

    Mahsa Derakhshan, Negin Golrezaei, Vahideh Manshadi, and Vahab Mirrokni. Product rank- ing on online platforms.���������� �������, 68(6):4024–4041, 2022. (Cited on page 8)

  4. [12]

    Constrained assortment optimiza- tion under the markov chain–based choice model.���������� �������, 66(2):698–721, 2020

    Antoine D ´esir, Vineet Goyal, Danny Segev, and Chun Ye. Constrained assortment optimiza- tion under the markov chain–based choice model.���������� �������, 66(2):698–721, 2020. (Cited on page 8)

  5. [13]

    Ai agents are changing how people shop

    Jur Gaarlandt, Wesley Korver, Nathan Furr, and Andrew Shipilov. Ai agents are changing how people shop. here’s what that means for brands.������� �������� ������, 26, 2025. (Cited on pages 2 and 7)

  6. [14]

    Examining the impact of ranking on consumer behavior and search engine revenue.���������� �������, 60(7):1632–1654, 2014

    Anindya Ghose, Panagiotis G Ipeirotis, and Beibei Li. Examining the impact of ranking on consumer behavior and search engine revenue.���������� �������, 60(7):1632–1654, 2014. (Cited on pages 8 and 30)

  7. [15]

    Frontiers: Can large language models capture human prefer- ences?��������� �������, 43(4):709–722, 2024

    Ali Goli and Amandeep Singh. Frontiers: Can large language models capture human prefer- ences?��������� �������, 43(4):709–722, 2024. (Cited on pages 9 and 30)

  8. [16]

    Near-optimal algorithms for the assortment planning problem under dynamic substitution and stochastic demand.���������� ��������, 64(1):219–235, 2016

    Vineet Goyal, Retsef Levi, and Danny Segev. Near-optimal algorithms for the assortment planning problem under dynamic substitution and stochastic demand.���������� ��������, 64(1):219–235, 2016. (Cited on page 8)

  9. [17]

    Design- ing algorithmic delegates: The role of indistinguishability in human–ai handoff

    Sophie Greenwood, Karen Levy, Solon Barocas, Hoda Heidari, and Jon Kleinberg. Design- ing algorithmic delegates: The role of indistinguishability in human–ai handoff. In�������� 32 ���� �� ��� ���� ��� ���������� �� ��������� ��� ����������� ����, 2025. Also available as arXiv...

  10. [18]

    Delegating to multiple agents

    MohammadTaghi Hajiaghayi, Keivan Rezaei, and Suho Shin. Delegating to multiple agents. In����������� �� ��� ���� ��� ���������� �� ��������� ��� ����������� ����, pages 861–885,

  11. [19]

    Lilium: ebay’s large language models for e-commerce.����� �������� ����������������, 2024

    Christian Herold, Michael Kozielski, Leonid Ekimov, Pavel Petrushkov, Pierre-Yves Vanden- bussche, and Shahram Khadivi. Lilium: ebay’s large language models for e-commerce.����� �������� ����������������, 2024. (Cited on page 8)

  12. [20]

    Large language models as simulated economic agents: What can we learn from homo silicus? Technical report, National Bureau of Economic Research, 2023

    John J Horton. Large language models as simulated economic agents: What can we learn from homo silicus? Technical report, National Bureau of Economic Research, 2023. (Cited on page 9)

  13. [21]

    The dawn of gui agent: A preliminary case study with claude 3.5 computer use.���� �������� ����������������, 2024

    Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. The dawn of gui agent: A preliminary case study with claude 3.5 computer use.���� �������� ����������������, 2024. (Cited on pages 37 and 38)

  14. [22]

    Social status and badge design

    Nicole Immorlica, Greg Stoddard, and Vasilis Syrgkanis. Social status and badge design. ����� �������� ���������������, 2013. (Cited on page 8)

  15. [23]

    Shopping mmlu: A massive multi-task online shopping benchmark for large language models.����� �������� ����������������, 2024

    Yilun Jin, Zheng Li, Chenwei Zhang, Tianyu Cao, Yifan Gao, Pratik Jayarao, Mao Li, Xin Liu, Ritesh Sarkhel, Xianfeng Tang, et al. Shopping mmlu: A massive multi-task online shopping benchmark for large language models.����� �������� ����������������, 2024. (Cited on page 8)

  16. [24]

    The cost of dynamic reasoning: De- mystifying ai agents and test-time scaling from an ai infrastructure perspective.���� �������� ������������������, 2025

    Jiin Kim, Byeongjun Shin, Jinha Chung, and Minsoo Rhu. The cost of dynamic reasoning: De- mystifying ai agents and test-time scaling from an ai infrastructure perspective.���� �������� ������������������, 2025. (Cited on page 39)

  17. [25]

    Delegated search approximates efficient search

    Jon Kleinberg and Robert Kleinberg. Delegated search approximates efficient search. In ����������� �� ��� ���� ��� ���������� �� ��������� ��� ����������� ����, pages 287–302,

  18. [26]

    Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.����� �������� ����������������,

    Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.����� �������� ����������������,

  19. [27]

    Assortment planning: Re- view of literature and industry practice.������ ������ ����� ����������, pages 209–246, 2008

    A G ¨urhan K ¨ok, Marshall L Fisher, and Ramnath Vaidyanathan. Assortment planning: Re- view of literature and industry practice.������ ������ ����� ����������, pages 209–246, 2008. (Cited on page 8)

  20. [28]

    Harnessing natural experiments to quantify the causal effect of badges.����� �������� ����������������, 2017

    Tomasz Kusmierczyk and Manuel Gomez-Rodriguez. Harnessing natural experiments to quantify the causal effect of badges.����� �������� ����������������, 2017. (Cited on page 8)

  21. [29]

    Product badges and consumer choice on digital platforms.��������� �� ���� �������, 2024

    Markus Lill, Nastasia Gallitz, Lucas Stich, and Martin Spann. Product badges and consumer choice on digital platforms.��������� �� ���� �������, 2024. (Cited on pages 7, 8, 18, and 30) 33

  22. [30]

    Deepshop: A benchmark for deep research shopping agents.����� �������� ����������������, 2025

    Yougang Lyu, Xiaoyu Zhang, Lingyong Yan, Maarten de Rijke, Zhaochun Ren, and Xiuy- ing Chen. Deepshop: A benchmark for deep research shopping agents.����� �������� ����������������, 2025. (Cited on pages 8 and 37)

  23. [31]

    Androidworld: A dynamic benchmarking environment for autonomous agents

    Arnav Matiana, Vasu Agarwal, Amol Naik, Shannon Zhang, Karthik Tirumala, Qixing Huang, et al. Androidworld: A dynamic benchmarking environment for autonomous agents. ����� �������� ����������������, 2024. (Cited on page 8)

  24. [32]

    Conditional logit analysis of qualitative choice behavior

    Daniel McFadden. Conditional logit analysis of qualitative choice behavior. In Paul Zarem- bka, editor,��������� �� ������������, pages 105–142. Academic Press, 1974. (Cited on pages 4 and 16)

  25. [33]

    Patil, Ion Stoica, and Joseph E

    Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. Memgpt: Towards llms as operating systems.����� �������� ����������������, 2024. (Cited on page 37)

  26. [34]

    O’Brien, Carrie J

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In�������� ���� �� ��� ���� ������ ��� ��������� �� ���� ��������� �������� ��� ���������� ����� ����, p...

  27. [35]

    ecellm: Generalizing large lan- guage models for e-commerce from large-scale, high-quality instruction data.����� �������� ����������������, 2024

    Bo Peng, Xinyi Ling, Ziru Chen, Huan Sun, and Xia Ning. ecellm: Generalizing large lan- guage models for e-commerce from large-scale, high-quality instruction data.����� �������� ����������������, 2024. (Cited on page 8)

  28. [36]

    A mega- study of digital twins reveals strengths, weaknesses and opportunities for further improve- ment.����� �������� ����������������, 2025

    Tiany Peng, George Gui, Daniel J Merlau, Grace Jiarui Fan, Malek Ben Sliman, Melanie Brucks, Eric J Johnson, Vicki Morwitz, Abdullah Althenayyan, Silvia Bellezza, et al. A mega- study of digital twins reveals strengths, weaknesses and opportunities for further improve- ment.��...

  29. [37]

    A survey of efficient rea- soning for large reasoning models: Language, multimodality, and beyond.����� �������� ����������������, March 2025

    Xiaoye Qu, Yafu Li, Zhaochen Su, Weigao Sun, Jianhao Yan, Dongrui Liu, Ganqu Cui, Daizong Liu, Shuxian Liang, Junxian He, Peng Li, Wei Wei, Jing Shao, Chaochao Lu, Yue Zhang, Xian-Sheng Hua, Bowen Zhou, and Yu Cheng. A survey of efficient rea- soning for large reasoning models...

  30. [38]

    Springer US, 2022

    Francesco Ricci, Lior Rokach, and Bracha Shapira.����������� ������� ��������. Springer US, 2022. (Cited on page 9)

  31. [39]

    Roumeliotis, and Manoj Karkee

    Ranjan Sapkota, Konstantinos I. Roumeliotis, and Manoj Karkee. Ai agents vs. agentic ai: A conceptual taxonomy, applications and challenges.����� �������� ����������������, 2025. (Cited on page 37)

  32. [40]

    Database report: Twin-2k-500: A data set for building digital twins of over 2,000 people 34 based on their answers to over 500 questions.��������� �������, 0(0), 2025

    Olivier Toubia, George Z Gui, Tianyi Peng, Daniel J Merlau, Ang Li, and Haozhe Chen. Database report: Twin-2k-500: A data set for building digital twins of over 2,000 people 34 based on their answers to over 500 questions.��������� �������, 0(0), 2025. (Cited on pages 9 and 30)

  33. [41]

    The power of rankings: Quantifying the effect of rankings on online con- sumer search and purchase decisions.��������� �������, 37(4):530–552, 2018

    Raluca M Ursu. The power of rankings: Quantifying the effect of rankings on online con- sumer search and purchase decisions.��������� �������, 37(4):530–552, 2018. (Cited on pages 7, 8, and 30)

  34. [42]

    Large language models for market re- search: A data-augmentation approach.����� �������� ����������������, 2024

    Mengxin Wang, Dennis J Zhang, and Heng Zhang. Large language models for market re- search: A data-augmentation approach.����� �������� ����������������, 2024. (Cited on page 9)

  35. [43]

    Agent.xpu: Efficient scheduling of agentic llm workloads on heterogeneous soc

    Xinming Wei, Jiahao Zhang, Haoran Li, Jiayu Chen, Rui Qu, Maoliang Li, Xiang Chen, and Guojie Luo. Agent.xpu: Efficient scheduling of agentic llm workloads on heterogeneous soc. ����� �������� ����������������, June 2025. (Cited on page 39)

  36. [44]

    Osworld: Benchmark- ing multimodal agents for open-ended tasks in real computer environments

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. Osworld: Benchmark- ing multimodal agents for open-ended tasks in real computer environments. In�������� �� ������ ����������� �����...

  37. [45]

    Pumgpt: A large vision-language model for product understanding.����� �������� ����������������, 2023

    Wei Xue, Zongyi Guo, Baoliang Cui, Zheng Xing, Xiaoyi Zeng, Xiufei Wang, Shuhui Wu, and Weiming Lu. Pumgpt: A large vision-language model for product understanding.����� �������� ����������������, 2023. (Cited on page 8)

  38. [46]

    Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press

    John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated soft- ware engineering.����� �������� ����������������, 2024. (Cited on page 8)

  39. [47]

    Webshop: Towards scal- able real-world web interaction with grounded language agents

    Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scal- able real-world web interaction with grounded language agents. In�������� �� ������ ������ ������ ���������� �������, volume 35, pages 20744–20757, 2022. (Cited on pages 8 and 37)

  40. [48]

    Ui-tars: Pioneering auto- mated gui interaction with native agents.����� �������� ����������������, 2025

    Ziyu Zhao, Zhenguang Li, Xiaoqin Pan, Jiale Chen, and Shu Li. Ui-tars: Pioneering auto- mated gui interaction with native agents.����� �������� ����������������, 2025. (Cited on page 8)

  41. [49]

    Gpt-4v(ision) is a generalist web agent, if grounded.����� �������� ����������������, 2024

    Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. Gpt-4v(ision) is a generalist web agent, if grounded.����� �������� ����������������, 2024. (Cited on page 7)

  42. [50]

    Webarena: A realistic web environment for building autonomous agents.����� �������� ����������������, 2024

    Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents.����� �������� ����������������, 2024. (Cited on page 7)

  43. [51]

    agents” are simply a VLM (or LLM) given the ability to interact with the external world through “tools

    Thomas P Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong. 35 Personalllm: Tailoring llms to individual preferences.����� �������� ����������������, 2024. (Cited on page 9) 36 Appendix Organization of the Appendix.In Section A, we provide further details ...

  44. [52]

    Navigates to the target platform

  45. [53]

    Searches for a product

  46. [54]

    Captures a screenshot

  47. [55]

    batch processing

    Selects a product using the provided tool ����seeks to collapse an agent’s journey to the linear sequence stated above, while accurately emulating a production-ready shopping agent, interacting with the underlying VLM models in a uniform manner, and ensuring that the choice be...

  48. [56]

    prioritize price

    We can use the estimates in Table 2, (5.1) and (5.2) to understand how the different models trade-off different attributes. For example, we can be interested in understanding how the choice 45 ��� ��� ������ ��� ������� �� ��� �������� ��� ��������� ����������� ���� ��������� ...

  49. [60]

    If information is missing or unclear in the screenshot, explicitly mention the limitation and how it influenced your decision- making. �/instructions� ������ ��� ������� ������ ��� �� ����� ����� ���� �������� �� ������ ��� �������� �� ��� ������� You are a personal shopping a...

  50. [61]

    Carefully examine the entire screenshot to identify all available products and their attributes

  51. [62]

    Use the�add to cart�function when you are ready to buy a product

  52. [63]

    Before making your selection, explain your reasoning for choosing this product, including what factors influenced your decision and any assumptions you made about what would be best: - Your primary decision criteria and why you prioritized them - How each available product per...

  53. [64]

    selected product number

    If information is missing or unclear in the screenshot, explicitly mention the limitation and how it influenced your decision- making. �/instructions� ������ ��� ���������� ������ ��� �� ����� ����� ���� �������� �� ������ ��� �������� �� ��� ������� ��� ���������� ����� E.2 P...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.