REVIEW 1 major objections 1 minor 22 references
Well-known brands receive every LLM recommendation for equivalent products, but this monopoly ends with a competitor's sub-0.1-star rating edge.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-27 01:37 UTC pith:CQO4R2JH
load-bearing objection The experiments show a conditional monopoly for established brands in LLM product recs when specs match, broken by tiny rating edges or marketing language, but scope is narrow. the 1 major comments →
Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
When products share identical specifications, LLMs recommend well-known brands 100 percent of the time, producing an Incumbent Advantage Index of 10.0; the monopoly dissolves when any competitor holds a rating advantage below 0.1 stars, authority-style marketing language disrupts the pattern at a bias surplus value of 0.17 rating points, and simultaneous adoption of optimization strategies by all brands reduces individual payoffs from 0.802 to 0.007 while leaving non-participants with zero recommendations.
What carries the argument
The Incumbent Advantage Index (IAI), a measure that records the share of recommendations captured by the leading brand when product attributes are held constant.
Load-bearing premise
Skincare products stand in for all experience goods whose quality buyers must infer from brand reputation, and the three tested LLMs represent typical commercial models in recommendation tasks.
What would settle it
Repeat the product queries on the same models after assigning a 0.05-star rating advantage to a lesser-known brand and check whether recommendations remain 100 percent for the incumbent.
If this is right
- A competitor overcomes the monopoly with less than a 0.1-star rating advantage.
- Authority-style marketing language delivers an advantage equivalent to a 0.17-star rating gain.
- When every brand applies the same optimization approach, individual payoffs fall sharply and non-adopters receive no recommendations.
- The bias pattern holds for both experience goods and search goods in the robustness check.
Where Pith is reading between the lines
- Platforms may need explicit mechanisms to prevent brand lock-in from becoming the default outcome of LLM queries.
- Smaller brands lacking resources for optimization could face permanent exclusion from visibility.
- The observed social dilemma suggests that widespread generative engine optimization could function as a prisoner's dilemma for marketers.
- Regulators might examine how generative engines reshape competition in the same way search rankings have been scrutinized.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper examines brand bias in LLM product recommendations through experiments on skincare products (experience goods) across three commercial LLMs (GPT-4o-mini, Claude Sonnet, Gemini 3 Flash), with a robustness check on search goods. Key findings include: (1) a Conditional Monopoly in which well-known brands receive 100% of recommendations (IAI = 10.0) when all products have identical specifications, but this dominance vanishes with a competitor's rating advantage of less than +0.1 stars; (2) authority-style marketing language (including fabricated clinical claims) breaks the monopoly at a Bias Surplus Value equivalent to +0.17 rating points, with model-specific responses; and (3) a social dilemma in multi-brand GEO competition where universal adoption reduces individual payoffs from +0.802 to +0.007 in the payoff proxy, with non-participants receiving zero recommendations.
Significance. If the reported thresholds and collective-action effects hold, the work identifies concrete mechanisms by which brand reputation and optimization tactics shape LLM-mediated markets, with implications for competition policy and the study of generative engine optimization as both a security issue and a marketing practice. The multi-model design and category robustness check provide an empirical foundation that could be extended to quantify divergence across LLMs and product types.
major comments (1)
- [Experiments (including robustness check)] The headline claims concern behavior in 'LLM recommendation systems' broadly, yet rest on results from only three models and primarily the skincare category. The abstract notes a robustness check on search goods but provides no quantitative comparison of effect sizes, model divergence, or category-specific differences; without such data, the extension from the tested cases to the stated domain remains unsupported.
minor comments (1)
- [Abstract] The abstract summarizes results but omits essential methodological details (sample sizes, statistical tests, controls, exact prompt templates) that would allow readers to assess the experiments at a glance; these should be added even if elaborated in the main text.
Simulated Author's Rebuttal
We thank the referee for the detailed and constructive report. We address the major comment on experimental scope and generalizability below, and commit to revisions that strengthen the manuscript without overstating current evidence.
read point-by-point responses
-
Referee: [Experiments (including robustness check)] The headline claims concern behavior in 'LLM recommendation systems' broadly, yet rest on results from only three models and primarily the skincare category. The abstract notes a robustness check on search goods but provides no quantitative comparison of effect sizes, model divergence, or category-specific differences; without such data, the extension from the tested cases to the stated domain remains unsupported.
Authors: We agree that the current manuscript does not provide sufficient quantitative detail on the robustness check to fully support broad claims. The three models tested are the leading commercial LLMs available during the study period, and skincare was selected as the primary category because it exemplifies experience goods where brand reputation is salient; the search-goods check was intended only as an initial robustness probe. In revision we will add a dedicated subsection (with tables) reporting direct comparisons of key metrics—IAI scores, monopoly rates, and Bias Surplus Values—between skincare and search goods, plus model-by-model effect sizes and divergence statistics. This will allow readers to assess the degree of consistency. We will also qualify the abstract and introduction to frame the results as initial evidence from leading models and two product types rather than a comprehensive survey of all LLM recommendation systems. revision: yes
Circularity Check
No circularity: empirical results from direct LLM experiments
full rationale
The paper reports controlled experiments measuring LLM recommendation frequencies for skincare products across three models. Central claims (Conditional Monopoly with IAI=10.0, bias surplus thresholds, social dilemma payoffs) are direct observational outputs from prompt-based tests, not derived via equations, fitted parameters, or self-citation chains. No load-bearing steps reduce to inputs by construction; the work is self-contained against external benchmarks.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption Skincare products represent experience goods where consumers rely on brand reputation.
read the original abstract
Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel. We study brand dynamics in LLM recommendations using skincare products -- a category where consumers cannot easily judge quality before buying and must rely on brand reputation -- across three commercial LLMs (GPT-4o-mini, Claude Sonnet, Gemini 3 Flash), with a robustness check on search goods. In three experiments, we find: (1) a Conditional Monopoly where well-known brands get recommended 100% of the time (IAI = 10.0) when all products have the same specifications, but this dominance disappears with less than a +0.1-star rating advantage for a competitor; (2) authority-style marketing language, including fabricated clinical-evidence claims, breaks this monopoly at a Bias Surplus Value equal to +0.17 rating points, with each model responding differently; and (3) a social dilemma in multi-brand GEO competition: when all brands adopt the same optimization strategy, individual payoff falls from +0.802 to +0.007 in our payoff proxy, and non-participating brands receive zero recommendations in our tests. Our results suggest that generative engine optimization (GEO) should be studied not only as a security risk, but also as an emerging marketing practice that shapes market competition.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande. 2024. GEO : Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), pages 50--61
2024
-
[3]
Bagga, Yuhang Wu, Pranjal Aggarwal, and Ameet Deshpande
Puneet S. Bagga, Yuhang Wu, Pranjal Aggarwal, and Ameet Deshpande. 2025. E-GEO : A testbed for generative engine optimization in e-commerce. ArXiv preprint arXiv:2511.20867
work page internal anchor Pith review arXiv 2025
- [4]
-
[5]
Jessica Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, and Zhouhan Chen. 2024. Cognitive bias in decision-making with LLMs . In Findings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 12594--12613
2024
-
[6]
Giorgos Filandrianos, Angeliki Dimitriou, Maria Lymperaiou, Konstantinos Thomas, and Giorgos Stamou. 2025. Bias beware: The impact of cognitive biases on LLM -driven product recommendations. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP)
2025
-
[7]
Garrett Hardin. 1968. The tragedy of the commons. Science, 162(3859):1243--1248
1968
-
[8]
Mahammed Kamruzzaman, Hieu Minh Nguyen, and Gene Louis Kim. 2024. `` Global is good, local is bad?'': Understanding brand bias in LLMs . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 12704--12721
2024
-
[9]
Kevin Lane Keller. 1993. Conceptualizing, measuring, and managing customer-based brand equity. Journal of Marketing, 57(1):1--22
1993
- [10]
- [11]
-
[12]
Weiran Lin, Yizhong Wang, Lukas Bauer, Yun He, Minjoon Seo, and Daniel Khashabi. 2025. LLM whisperer: An inconspicuous attack to bias LLM responses. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI '25). ACM. Article 188375
2025
-
[13]
Fatemeh Nazary, Yashar Deldjoo, and Tommaso Di Noia. 2025. Poison- RAG : Adversarial data poisoning attacks on retrieval-augmented generation in recommender systems. In Advances in Information Retrieval (ECIR 2025), pages 239--254. Springer
2025
-
[14]
Phillip Nelson. 1970. Information and consumer behavior. Journal of Political Economy, 78(2):311--329
1970
-
[15]
Fredrik Nestaas, Edoardo Debenedetti, and Florian Tram \`e r. 2025. Adversarial search engine optimization for large language models. In Proceedings of the International Conference on Learning Representations (ICLR)
2025
-
[16]
OpenAI . 2026. Testing ads in ChatGPT . https://openai.com/index/testing-ads-in-chatgpt/. Published February 9, 2026. Accessed: 2026-05-03
2026
-
[17]
Samuel Pfrommer, Yuze Bai, Tanmay Gautam, and Somayeh Sojoudi. 2024. Ranking manipulation for conversational search engines. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9520--9534
2024
-
[18]
Talboy and Elizabeth Fuller
Andr \'e s N. Talboy and Elizabeth Fuller. 2024. Challenging the status quo: Human bias in AI models? anchoring effects and mitigation strategies in large language models. Expert Systems with Applications, 255:124803
2024
-
[19]
Lanling Xu, Junjie Zhang, Bingqian Yu, Jie Zhang, Hongzhi Li, Jingjing Li, and Xin Zhao. 2025. A survey on LLM -powered agents for recommender systems. In Findings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8015--8049
2025
- [20]
- [21]
-
[22]
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2025. PoisonedRAG : Knowledge corruption attacks to retrieval-augmented generation of large language models. In Proceedings of the 34th USENIX Security Symposium
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.