REVIEW 4 major objections 5 minor 61 references
Actionable Counterfactual Explanations Using Bayesian Networks and Path Planning with Applications to Environmental Quality Improvement
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Frontier AI shopping agents exhibit large position biases that persist across interfaces and invert when models are upgraded.
desk verdict The body under review is a solid, carefully randomized measurement study of AI shopping agents, but it is not the paper the arXiv metadata advertises, and its headline claims outrun the evidence by skipping a human baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
ACES (Agentic e-Commerce Simulator): a provider-agnostic evaluation sandbox that pairs a vision-language agent, controlling a browser through tools, with a programmable mock storefront rendering an eight-product grid in a randomized layout. Its engine is the randomized controlled trial: positions, tag assignments, prices, ratings, and review counts are exogenously varied, and a conditional logit model recovers causal choice probabilities and price-equivalent trade-offs between levers. The same trials are reproduced in a headless variant that exposes only a ranked JSON list with no images, which separates visual parsing artifacts from deep-seated model priors.
What would settle it
Run the same position-randomization trials with an agent that browses multiple pages, opens product-detail pages, and carries user context; if the position bias and sponsored-tag penalty shrink to near zero under full-funnel browsing, the claim that these distortions are structural features of agentic commerce would fail. A second test: a released model generation that preserves its predecessor's position preferences would falsify the claim that upgrades systematically invert or reshuffle these biases.
Extended reading notes
Core claim
The paper's central claim is that frontier AI shopping agents choose products in ways that are simultaneously homogeneous, unstable, and causally responsive to platform levers. Across hundreds of randomized trials per product category, selection shares collapse onto a few modal products while other brands are never chosen; model upgrades reshuffle shares drastically, with the Fitbit Inspire's share jumping from 25% to 77% after one Claude upgrade and falling from 25% to 6% after a GPT upgrade. Conditional-logit estimates show position effects large enough that moving a product from the bottom-right slot to the top row can raise its selection rate five-fold, and the preferred position of GPT-4.1 is the least preferred position of its successor GPT-5.1. Because badge assignment and position were randomized, the paper reads the sponsored-tag penalty and the Overall Pick lift as causal, with the endorsement worth as much as a 65-138% price increase in price-equivalent terms. A seller-side agent making one-shot, query-conditional description edits produced significant share gains in five of six buyer models, with office-lamp gains of up to 80.4 percentage points from front-loading the word 'Office'.
Load-bearing premise
The load-bearing premise is that a one-shot choice by an unassisted vision-language agent over an eight-product grid, from a generic prompt in a synthetic mock store, faithfully approximates real agentic shopping, with sampled selection frequencies treated as market shares.
Editorial extensions
If this is right
- If AI agents mediate a growing share of purchases, product rankings become a first-order market lever, and there is no universal 'top' slot because the best position is model-specific.
- Platform endorsements like Overall Pick act as powerful credibility signals to agents while sponsored tags carry a credibility penalty, shifting how advertising and promotion budgets should be valued.
- Model updates function as exogenous demand shocks, so sellers with static listings can gain or lose market share overnight even when catalogs are unchanged.
- Simple seller-side text optimization can unlock large share gains, implying that listing titles will become a strategic battlefield and that static descriptions are suboptimal against algorithmic buyers.
- Because position effects persist in headless/API settings and resist 'ignore position' prompting, the distortions are unlikely to disappear without deliberate design intervention.
- The paper's finding that agents concentrate choice on a few modal products implies concentration risk: dominant agents could suppress niche brands that would compete normally under human demand.
Reading between the lines
- A testable extension is to run the same position-randomization trials with agents that browse multiple pages, open product-detail pages, and carry user context; if position effects attenuate under full-funnel browsing, the 'deep-seated characteristic' reading would need to be revised downward.
- The price-equivalent trade-offs the paper computes suggest that platforms could price endorsements, premium slots, and even listing-text services in an 'agent currency', a design implication the paper states only implicitly.
- The inverted position preferences between adjacent model generations hint that post-training choices, rather than the pretrained knowledge base, drive much of the bias; direct attribution experiments on model checkpoints could test this.
- If choice homogeneity persists as agents scale, marketplaces may need diversity-promoting mechanisms or neutrality audits, since AI-mediated demand could otherwise produce winner-take-all outcomes regardless of underlying consumer taste.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript (arXiv:2508.02630v3, as appears in the full text) introduces ACES, a controllable agent-platform sandbox for auditing AI shopping agents. Using randomized experiments with frontier vision-language models (Claude, GPT, Gemini) in an eight-product grid, the paper measures instruction following, market shares, and causal responses to position, badges, price, ratings, and reviews via conditional logit models. It reports strong model-dependent position biases that persist in headless/API settings, a causal penalty for Sponsored tags, a large positive effect of Overall Pick endorsements, and significant seller-side market-share gains from AI-generated title edits. The paper argues that agentic markets are volatile, model-dependent, and fundamentally different from human-centric commerce, and proposes continuous auditing frameworks. The core empirical design is sound: positions, badges, and attributes are randomized, and the seller experiments reuse identical product shuffles pre/post, enabling causal attribution. However, several claims are overstated relative to the evidence, and the submitted abstract does not match the manuscript's content.
Significance. If the findings hold, this paper is significant for platform design, seller strategy, and AI governance. It provides one of the first rigorous, provider-agnostic frameworks for causally auditing AI shopping agents, with randomized identification of position and badge effects. The headless-interface and prompt-variation robustness checks strengthen the case that these biases are not artifacts of screenshot parsing. The paper ships code and data (GitHub, HuggingFace), supporting replication. The finding that simple title edits can swing market shares—with a concrete mechanism (keyword front-loading)—is a practical and falsifiable contribution. However, the lack of a human baseline and the strong cross-model comparisons limit the external-validity claims.
major comments (4)
- [Abstract (submitted text)] The abstract at the beginning of the manuscript describes a different paper—actionable counterfactual explanations using Bayesian networks and an EPA dataset—and does not correspond to the full text, which is about AI shopping agents in ACES. This mismatch must be corrected; it is not a minor typo and would mislead readers. The editor should verify that the correct abstract is attached.
- [Abstract and Section 4] The claim that agentic markets are 'fundamentally different from human-centric commerce' is not supported by a direct human baseline. The paper defers a human-subject study (Section 8) but states the stronger claim in the abstract. Existing human literature (e.g., Ursu 2018, Lill et al. 2024) shows similar position and badge effects, so without running the same grid and prompt with human participants, the 'fundamentally different' assertion is an interpretation rather than a measured result. Please soften the claim or add a human baseline.
- [Section 4] Market-share figures (Figure 1 and Figures 5–6) report point estimates from 200 trials per category without standard errors or confidence intervals. Claims such as 'Fitbit Inspire jumped from 45% to 77%' need uncertainty quantification; with n=200, the standard error of a proportion is roughly 3.5 percentage points, so formal cross-model hypothesis tests should be reported to support the 'drastic swings' language.
- [Section 6.2] The seller agent is given the exact trial-level sales data of the buying agent and full competitor listings, an information advantage that real sellers would not have. The paper should state explicitly that the reported market-share gains are upper bounds under an informed seller, and discuss how the effects might attenuate with realistic, noisier information. This caveat is central to the 'AI-SEO can drive significant gains' contribution.
minor comments (5)
- [Section 5.2, Table 2 note] The footnote states that products never selected were excluded from the conditional logit dataset, reducing the average alternatives per choice set to 6.8. This is standard, but the paper should clarify that this does not affect the consistency of the estimates and should report how many products were dropped per model.
- [References] In Section 1, the text cites '[13]' for 'race and gender disparities in AI-based hiring', but reference [13] is Gaarlandt et al., 'AI agents are changing how people shop'. This citation appears to point to the wrong source; please verify and correct.
- [Section 7.1] The temporal drift analysis re-ran experiments in September 2025 with the same model names (e.g., Claude Sonnet 4). The paper should clarify whether these are the exact same model checkpoints or whether API updates could have occurred, since model finalization is a studied phenomenon in the paper itself.
- [Section 5.2] The conditional logit specification (5.1) assumes IIA; the paper does not discuss this assumption or test its plausibility in the product-choice context. A brief discussion or a robustness check (e.g., nested logit or mixed logit) would strengthen the interpretation of the coefficients.
- [Table 11] The price-equivalent trade-off for Gemini 3.0 Pro Preview reports a Row 1 premium of +159.7% and an Overall Pick premium of +159.3%. These values are much larger than for other models; please verify the arithmetic and the standard errors.
Circularity Check
No circularity: an empirical measurement paper whose headline quantities are direct experimental outcomes, not derived predictions or renamed inputs.
full rationale
The paper does not derive its headline claims from a fitted model, a self-citation, or a uniqueness theorem. Market shares are direct selection frequencies from randomized trials (Section 4.1: 'we compute the induced market shares, representing the selection rates of different products across the 400 experiments'). The conditional logit estimates in Section 5.2 are descriptive summaries of the same experimental choices, and the position-probability plots are explicitly labeled 'model-based estimates' for illustration, not independent predictions. Seller-side gains are measured as observed pre/post differences in selection probability under identical shuffles (Section 6.2: 'we reuse the identical product shuffles in the pre/post modification experiments'), so the outcome is experimental rather than derived from the fitted model. The headless interface and prompt-robustness checks are separate runs with new choices. No load-bearing premise is justified by a self-citation, and no prediction is a renamed experimental input. The absence of a human baseline is an external-validity limitation, not a circularity.
Assumptions & free parameters
free parameters (6)
- Conditional logit price coefficient =
estimated per model; Tables 2-3, 12-19
- Conditional logit rating coefficient =
estimated per model; Tables 2-3, 12-19
- Conditional logit number-of-reviews coefficient =
estimated per model; Tables 2-3, 12-19
- Position dummies (Row 1, Columns 1-3, or Positions 1-7 in headless) =
estimated per model; Tables 2-3, 17-19
- Badge dummies (Sponsored, Overall Pick, Scarcity) =
estimated per model; Tables 2-3
- Product fixed effects =
estimated per product per model
assumptions (4)
- domain assumption Conditional logit with linear index and IIA (Eq. 5.1-5.2) captures agent choice behavior
- domain assumption One-shot choice over an eight-product grid with a generic prompt approximates real agentic shopping (Section 2)
- domain assumption Model configuration (temperature 1.0, minimal reasoning budgets) is a valid deployment baseline (Table 5)
- domain assumption Author-selected eight-product assortments from Amazon listings are representative (Section 2)
invented entities (1)
-
ACES (Agentic e-Commerce Simulator)
independent evidence
Cite this review
Pith. "Pith review of Actionable Counterfactual Explanations Using Bayesian Networks and Path Planning with Applications to Environmental Quality Improvement." pith.science (2026). https://pith.science/paper/4JIKVWK5
@misc{pith2026250802634,
author = {Pith},
title = {Pith review of: Actionable Counterfactual Explanations Using Bayesian Networks and Path Planning with Applications to Environmental Quality Improvement},
year = {2026},
howpublished = {\url{https://pith.science/paper/4JIKVWK5}},
note = {Machine review of arXiv:2508.02634}
}
read the original abstract
Counterfactual explanations study what should have changed in order to get an alternative result, enabling end-users to understand machine learning mechanisms with counterexamples. Actionability is defined as the ability to transform the original case to be explained into a counterfactual one. We develop a method for actionable counterfactual explanations that, unlike predecessors, does not directly leverage training data. Rather, data is only used to learn a density estimator, creating a search landscape in which to apply path planning algorithms to solve the problem and masking the endogenous data, which can be sensitive or private. We put special focus on estimating the data density using Bayesian networks, demonstrating how their enhanced interpretability is useful in high-stakes scenarios in which fairness is raising concern. Using a synthetic benchmark comprised of 15 datasets, our proposal finds more actionable and simpler counterfactuals than the current state-of-the-art algorithms. We also test our algorithm with a real-world Environmental Protection Agency dataset, facilitating a more efficient and equitable study of policies to improve the quality of life in United States of America counties. Our proposal captures the interaction of variables, ensuring equity in decisions, as policies to improve certain domains of study (air, water quality, etc.) can be detrimental in others. In particular, the sociodemographic domain is often involved, where we find important variables related to the ongoing housing crisis that can potentially have a severe negative impact on communities.
Reference graph
Works this paper leans on
-
[1]
Saaket Agashe, Kyle Wong, Vincent Tu, Jiachen Yang, Ang Li, and Xin Eric Wang. Agent s2: A compositional generalist-specialist framework for computer use agents.����� �������� ����������������, 2025. (Cited on pages 8 and 37)
work page 2025
-
[2]
A model of delegated project choice.������������, 78(1):213–244, 2010
Mark Armstrong and John Vickers. A model of delegated project choice.������������, 78(1):213–244, 2010. (Cited on page 9)
work page 2010
-
[3]
Archit Bansal, Kunal Banerjee, and Abhijnan Chakraborty. These deals won’t last! longevity, uniformity and bias in product badge assignment in e-commerce platforms.����� �������� ����������������, 2022. (Cited on page 8)
work page 2022
-
[4]
Tammo H.A. Bijmolt, Harald J. Van Heerde, and Rik G.M. Pieters. New empirical generaliza- tions on the determinants of price elasticity.������� �� ��������� ��������, 42(2):141–156, 2005. (Cited on page 18) 31
work page 2005
-
[5]
Win- dows agent arena: Evaluating multi-modal os agents at scale.����� �������� ����������������,
Rogerio Bonatti, David Zhao, Francesco Bonacci, David Dupont, Sasan Abdali, Yifan Li, Yix- uan Lu, Jignesh Wagle, Kazuhito Koishida, Andres Bucker, Leon Jang, and Zihang Hui. Win- dows agent arena: Evaluating multi-modal os agents at scale.����� �������� ����������������,
-
[6]
Using llms for market research.������� �������� ������ ��������� ���� ������� �����, (23-062), 2023
James Brand, Ayelet Israeli, and Donald Ngwe. Using llms for market research.������� �������� ������ ��������� ���� ������� �����, (23-062), 2023. (Cited on page 9)
work page 2023
-
[7]
Online search and product rankings: A double logit approach
Giovanni Compiani, Gregory Lewis, Sida Peng, and Peichun Wang. Online search and product rankings: A double logit approach. Technical report, Working paper, University of Chicago, Chicago, IL, 2022. (Cited on page 8)
work page 2022
-
[8]
A shopping agent for addressing subjective product needs
Preetam Prabhu Srikar Dammu, Omar Alonso, and Barbara Poblete. A shopping agent for addressing subjective product needs. In����������� �� ��� ���������� ��� ������������� ���� ������� �� ��� ������ ��� ���� ������, pages 1032–1035, 2025. (Cited on page 8)
work page 2025
Show all 61 references
-
[9]
Assortment optimization under variants of the nested logit model.���������� ��������, 62(2):250–273, 2014
James M Davis, Guillermo Gallego, and Huseyin Topaloglu. Assortment optimization under variants of the nested logit model.���������� ��������, 62(2):250–273, 2014. (Cited on page 8)
2014
-
[10]
Mind2web: Towards a generalist agent for the web.�������� �� ������ ����������� ���������� �������, 36:28091–28114, 2023
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web.�������� �� ������ ����������� ���������� �������, 36:28091–28114, 2023. (Cited on pages 7 and 31)
2023
-
[11]
Product rank- ing on online platforms.���������� �������, 68(6):4024–4041, 2022
Mahsa Derakhshan, Negin Golrezaei, Vahideh Manshadi, and Vahab Mirrokni. Product rank- ing on online platforms.���������� �������, 68(6):4024–4041, 2022. (Cited on page 8)
2022
-
[12]
Constrained assortment optimiza- tion under the markov chain–based choice model.���������� �������, 66(2):698–721, 2020
Antoine D ´esir, Vineet Goyal, Danny Segev, and Chun Ye. Constrained assortment optimiza- tion under the markov chain–based choice model.���������� �������, 66(2):698–721, 2020. (Cited on page 8)
2020
-
[13]
Ai agents are changing how people shop
Jur Gaarlandt, Wesley Korver, Nathan Furr, and Andrew Shipilov. Ai agents are changing how people shop. here’s what that means for brands.������� �������� ������, 26, 2025. (Cited on pages 2 and 7)
2025
-
[14]
Examining the impact of ranking on consumer behavior and search engine revenue.���������� �������, 60(7):1632–1654, 2014
Anindya Ghose, Panagiotis G Ipeirotis, and Beibei Li. Examining the impact of ranking on consumer behavior and search engine revenue.���������� �������, 60(7):1632–1654, 2014. (Cited on pages 8 and 30)
2014
-
[15]
Frontiers: Can large language models capture human prefer- ences?��������� �������, 43(4):709–722, 2024
Ali Goli and Amandeep Singh. Frontiers: Can large language models capture human prefer- ences?��������� �������, 43(4):709–722, 2024. (Cited on pages 9 and 30)
2024
-
[16]
Near-optimal algorithms for the assortment planning problem under dynamic substitution and stochastic demand.���������� ��������, 64(1):219–235, 2016
Vineet Goyal, Retsef Levi, and Danny Segev. Near-optimal algorithms for the assortment planning problem under dynamic substitution and stochastic demand.���������� ��������, 64(1):219–235, 2016. (Cited on page 8)
2016
-
[17]
Design- ing algorithmic delegates: The role of indistinguishability in human–ai handoff
Sophie Greenwood, Karen Levy, Solon Barocas, Hoda Heidari, and Jon Kleinberg. Design- ing algorithmic delegates: The role of indistinguishability in human–ai handoff. In�������� 32 ���� �� ��� ���� ��� ���������� �� ��������� ��� ����������� ����, 2025. Also available as arXiv...
2025 arXiv
-
[18]
Delegating to multiple agents
MohammadTaghi Hajiaghayi, Keivan Rezaei, and Suho Shin. Delegating to multiple agents. In����������� �� ��� ���� ��� ���������� �� ��������� ��� ����������� ����, pages 861–885,
-
[19]
Lilium: ebay’s large language models for e-commerce.����� �������� ����������������, 2024
Christian Herold, Michael Kozielski, Leonid Ekimov, Pavel Petrushkov, Pierre-Yves Vanden- bussche, and Shahram Khadivi. Lilium: ebay’s large language models for e-commerce.����� �������� ����������������, 2024. (Cited on page 8)
2024
-
[20]
Large language models as simulated economic agents: What can we learn from homo silicus? Technical report, National Bureau of Economic Research, 2023
John J Horton. Large language models as simulated economic agents: What can we learn from homo silicus? Technical report, National Bureau of Economic Research, 2023. (Cited on page 9)
2023
-
[21]
The dawn of gui agent: A preliminary case study with claude 3.5 computer use.���� �������� ����������������, 2024
Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. The dawn of gui agent: A preliminary case study with claude 3.5 computer use.���� �������� ����������������, 2024. (Cited on pages 37 and 38)
2024
-
[22]
Social status and badge design
Nicole Immorlica, Greg Stoddard, and Vasilis Syrgkanis. Social status and badge design. ����� �������� ���������������, 2013. (Cited on page 8)
2013
-
[23]
Shopping mmlu: A massive multi-task online shopping benchmark for large language models.����� �������� ����������������, 2024
Yilun Jin, Zheng Li, Chenwei Zhang, Tianyu Cao, Yifan Gao, Pratik Jayarao, Mao Li, Xin Liu, Ritesh Sarkhel, Xianfeng Tang, et al. Shopping mmlu: A massive multi-task online shopping benchmark for large language models.����� �������� ����������������, 2024. (Cited on page 8)
2024
-
[24]
The cost of dynamic reasoning: De- mystifying ai agents and test-time scaling from an ai infrastructure perspective.���� �������� ������������������, 2025
Jiin Kim, Byeongjun Shin, Jinha Chung, and Minsoo Rhu. The cost of dynamic reasoning: De- mystifying ai agents and test-time scaling from an ai infrastructure perspective.���� �������� ������������������, 2025. (Cited on page 39)
2025
-
[25]
Delegated search approximates efficient search
Jon Kleinberg and Robert Kleinberg. Delegated search approximates efficient search. In ����������� �� ��� ���� ��� ���������� �� ��������� ��� ����������� ����, pages 287–302,
-
[26]
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.����� �������� ����������������,
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.����� �������� ����������������,
-
[27]
Assortment planning: Re- view of literature and industry practice.������ ������ ����� ����������, pages 209–246, 2008
A G ¨urhan K ¨ok, Marshall L Fisher, and Ramnath Vaidyanathan. Assortment planning: Re- view of literature and industry practice.������ ������ ����� ����������, pages 209–246, 2008. (Cited on page 8)
2008
-
[28]
Harnessing natural experiments to quantify the causal effect of badges.����� �������� ����������������, 2017
Tomasz Kusmierczyk and Manuel Gomez-Rodriguez. Harnessing natural experiments to quantify the causal effect of badges.����� �������� ����������������, 2017. (Cited on page 8)
2017
-
[29]
Product badges and consumer choice on digital platforms.��������� �� ���� �������, 2024
Markus Lill, Nastasia Gallitz, Lucas Stich, and Martin Spann. Product badges and consumer choice on digital platforms.��������� �� ���� �������, 2024. (Cited on pages 7, 8, 18, and 30) 33
2024
-
[30]
Deepshop: A benchmark for deep research shopping agents.����� �������� ����������������, 2025
Yougang Lyu, Xiaoyu Zhang, Lingyong Yan, Maarten de Rijke, Zhaochun Ren, and Xiuy- ing Chen. Deepshop: A benchmark for deep research shopping agents.����� �������� ����������������, 2025. (Cited on pages 8 and 37)
2025
-
[31]
Androidworld: A dynamic benchmarking environment for autonomous agents
Arnav Matiana, Vasu Agarwal, Amol Naik, Shannon Zhang, Karthik Tirumala, Qixing Huang, et al. Androidworld: A dynamic benchmarking environment for autonomous agents. ����� �������� ����������������, 2024. (Cited on page 8)
2024
-
[32]
Conditional logit analysis of qualitative choice behavior
Daniel McFadden. Conditional logit analysis of qualitative choice behavior. In Paul Zarem- bka, editor,��������� �� ������������, pages 105–142. Academic Press, 1974. (Cited on pages 4 and 16)
1974
-
[33]
Patil, Ion Stoica, and Joseph E
Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. Memgpt: Towards llms as operating systems.����� �������� ����������������, 2024. (Cited on page 37)
2024
-
[34]
O’Brien, Carrie J
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In�������� ���� �� ��� ���� ������ ��� ��������� �� ���� ��������� �������� ��� ���������� ����� ����, p...
2023
-
[35]
ecellm: Generalizing large lan- guage models for e-commerce from large-scale, high-quality instruction data.����� �������� ����������������, 2024
Bo Peng, Xinyi Ling, Ziru Chen, Huan Sun, and Xia Ning. ecellm: Generalizing large lan- guage models for e-commerce from large-scale, high-quality instruction data.����� �������� ����������������, 2024. (Cited on page 8)
2024
-
[36]
A mega- study of digital twins reveals strengths, weaknesses and opportunities for further improve- ment.����� �������� ����������������, 2025
Tiany Peng, George Gui, Daniel J Merlau, Grace Jiarui Fan, Malek Ben Sliman, Melanie Brucks, Eric J Johnson, Vicki Morwitz, Abdullah Althenayyan, Silvia Bellezza, et al. A mega- study of digital twins reveals strengths, weaknesses and opportunities for further improve- ment.��...
2025
-
[37]
A survey of efficient rea- soning for large reasoning models: Language, multimodality, and beyond.����� �������� ����������������, March 2025
Xiaoye Qu, Yafu Li, Zhaochen Su, Weigao Sun, Jianhao Yan, Dongrui Liu, Ganqu Cui, Daizong Liu, Shuxian Liang, Junxian He, Peng Li, Wei Wei, Jing Shao, Chaochao Lu, Yue Zhang, Xian-Sheng Hua, Bowen Zhou, and Yu Cheng. A survey of efficient rea- soning for large reasoning models...
2025
-
[38]
Springer US, 2022
Francesco Ricci, Lior Rokach, and Bracha Shapira.����������� ������� ��������. Springer US, 2022. (Cited on page 9)
2022
-
[39]
Roumeliotis, and Manoj Karkee
Ranjan Sapkota, Konstantinos I. Roumeliotis, and Manoj Karkee. Ai agents vs. agentic ai: A conceptual taxonomy, applications and challenges.����� �������� ����������������, 2025. (Cited on page 37)
2025
-
[40]
Database report: Twin-2k-500: A data set for building digital twins of over 2,000 people 34 based on their answers to over 500 questions.��������� �������, 0(0), 2025
Olivier Toubia, George Z Gui, Tianyi Peng, Daniel J Merlau, Ang Li, and Haozhe Chen. Database report: Twin-2k-500: A data set for building digital twins of over 2,000 people 34 based on their answers to over 500 questions.��������� �������, 0(0), 2025. (Cited on pages 9 and 30)
2025
-
[41]
The power of rankings: Quantifying the effect of rankings on online con- sumer search and purchase decisions.��������� �������, 37(4):530–552, 2018
Raluca M Ursu. The power of rankings: Quantifying the effect of rankings on online con- sumer search and purchase decisions.��������� �������, 37(4):530–552, 2018. (Cited on pages 7, 8, and 30)
2018
-
[42]
Large language models for market re- search: A data-augmentation approach.����� �������� ����������������, 2024
Mengxin Wang, Dennis J Zhang, and Heng Zhang. Large language models for market re- search: A data-augmentation approach.����� �������� ����������������, 2024. (Cited on page 9)
2024
-
[43]
Agent.xpu: Efficient scheduling of agentic llm workloads on heterogeneous soc
Xinming Wei, Jiahao Zhang, Haoran Li, Jiayu Chen, Rui Qu, Maoliang Li, Xiang Chen, and Guojie Luo. Agent.xpu: Efficient scheduling of agentic llm workloads on heterogeneous soc. ����� �������� ����������������, June 2025. (Cited on page 39)
2025
-
[44]
Osworld: Benchmark- ing multimodal agents for open-ended tasks in real computer environments
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. Osworld: Benchmark- ing multimodal agents for open-ended tasks in real computer environments. In�������� �� ������ ����������� �����...
2024
-
[45]
Pumgpt: A large vision-language model for product understanding.����� �������� ����������������, 2023
Wei Xue, Zongyi Guo, Baoliang Cui, Zheng Xing, Xiaoyi Zeng, Xiufei Wang, Shuhui Wu, and Weiming Lu. Pumgpt: A large vision-language model for product understanding.����� �������� ����������������, 2023. (Cited on page 8)
2023
-
[46]
Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press
John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated soft- ware engineering.����� �������� ����������������, 2024. (Cited on page 8)
2024
-
[47]
Webshop: Towards scal- able real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scal- able real-world web interaction with grounded language agents. In�������� �� ������ ������ ������ ���������� �������, volume 35, pages 20744–20757, 2022. (Cited on pages 8 and 37)
2022
-
[48]
Ui-tars: Pioneering auto- mated gui interaction with native agents.����� �������� ����������������, 2025
Ziyu Zhao, Zhenguang Li, Xiaoqin Pan, Jiale Chen, and Shu Li. Ui-tars: Pioneering auto- mated gui interaction with native agents.����� �������� ����������������, 2025. (Cited on page 8)
2025
-
[49]
Gpt-4v(ision) is a generalist web agent, if grounded.����� �������� ����������������, 2024
Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su. Gpt-4v(ision) is a generalist web agent, if grounded.����� �������� ����������������, 2024. (Cited on page 7)
2024
-
[50]
Webarena: A realistic web environment for building autonomous agents.����� �������� ����������������, 2024
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents.����� �������� ����������������, 2024. (Cited on page 7)
2024
-
[51]
agents” are simply a VLM (or LLM) given the ability to interact with the external world through “tools
Thomas P Zollo, Andrew Wei Tung Siah, Naimeng Ye, Ang Li, and Hongseok Namkoong. 35 Personalllm: Tailoring llms to individual preferences.����� �������� ����������������, 2024. (Cited on page 9) 36 Appendix Organization of the Appendix.In Section A, we provide further details ...
2024
-
[52]
Navigates to the target platform
-
[53]
Searches for a product
-
[54]
Captures a screenshot
-
[55]
batch processing
Selects a product using the provided tool ����seeks to collapse an agent’s journey to the linear sequence stated above, while accurately emulating a production-ready shopping agent, interacting with the underlying VLM models in a uniform manner, and ensuring that the choice be...
2025
-
[56]
prioritize price
We can use the estimates in Table 2, (5.1) and (5.2) to understand how the different models trade-off different attributes. For example, we can be interested in understanding how the choice 45 ��� ��� ������ ��� ������� �� ��� �������� ��� ��������� ����������� ���� ��������� ...
-
[60]
If information is missing or unclear in the screenshot, explicitly mention the limitation and how it influenced your decision- making. �/instructions� ������ ��� ������� ������ ��� �� ����� ����� ���� �������� �� ������ ��� �������� �� ��� ������� You are a personal shopping a...
-
[61]
Carefully examine the entire screenshot to identify all available products and their attributes
-
[62]
Use the�add to cart�function when you are ready to buy a product
-
[63]
Before making your selection, explain your reasoning for choosing this product, including what factors influenced your decision and any assumptions you made about what would be best: - Your primary decision criteria and why you prioritized them - How each available product per...
-
[64]
selected product number
If information is missing or unclear in the screenshot, explicitly mention the limitation and how it influenced your decision- making. �/instructions� ������ ��� ���������� ������ ��� �� ����� ����� ���� �������� �� ������ ��� �������� �� ��� ������� ��� ���������� ����� E.2 P...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.