Pith. sign in

REVIEW 1 major objections 1 minor 22 references

Well-known brands receive every LLM recommendation for equivalent products, but this monopoly ends with a competitor's sub-0.1-star rating edge.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-27 01:37 UTC pith:CQO4R2JH

load-bearing objection The experiments show a conditional monopoly for established brands in LLM product recs when specs match, broken by tiny rating edges or marketing language, but scope is narrow. the 1 major comments →

arxiv 2606.17443 v1 pith:CQO4R2JH submitted 2026-06-16 cs.AI cs.CLcs.CY

Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems

classification cs.AI cs.CLcs.CY
keywords LLM recommendationsbrand biasincumbent advantagegenerative engine optimizationconditional monopolymarketing languagesocial dilemmarating advantage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tests how brands fare when consumers query LLMs for product suggestions in categories where quality cannot be inspected before purchase. It finds that identical products produce a conditional monopoly in which established brands capture 100 percent of recommendations, an effect quantified by an Incumbent Advantage Index of 10.0. This lock-in disappears once a rival gains even a modest rating lead or deploys authority-style claims, and it creates a social dilemma when every brand pursues the same optimization tactic. The results frame LLM recommendations as a new arena in which small quality signals and marketing language can rapidly reorder market access.

Core claim

When products share identical specifications, LLMs recommend well-known brands 100 percent of the time, producing an Incumbent Advantage Index of 10.0; the monopoly dissolves when any competitor holds a rating advantage below 0.1 stars, authority-style marketing language disrupts the pattern at a bias surplus value of 0.17 rating points, and simultaneous adoption of optimization strategies by all brands reduces individual payoffs from 0.802 to 0.007 while leaving non-participants with zero recommendations.

What carries the argument

The Incumbent Advantage Index (IAI), a measure that records the share of recommendations captured by the leading brand when product attributes are held constant.

Load-bearing premise

Skincare products stand in for all experience goods whose quality buyers must infer from brand reputation, and the three tested LLMs represent typical commercial models in recommendation tasks.

What would settle it

Repeat the product queries on the same models after assigning a 0.05-star rating advantage to a lesser-known brand and check whether recommendations remain 100 percent for the incumbent.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • A competitor overcomes the monopoly with less than a 0.1-star rating advantage.
  • Authority-style marketing language delivers an advantage equivalent to a 0.17-star rating gain.
  • When every brand applies the same optimization approach, individual payoffs fall sharply and non-adopters receive no recommendations.
  • The bias pattern holds for both experience goods and search goods in the robustness check.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Platforms may need explicit mechanisms to prevent brand lock-in from becoming the default outcome of LLM queries.
  • Smaller brands lacking resources for optimization could face permanent exclusion from visibility.
  • The observed social dilemma suggests that widespread generative engine optimization could function as a prisoner's dilemma for marketers.
  • Regulators might examine how generative engines reshape competition in the same way search rankings have been scrutinized.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. The paper examines brand bias in LLM product recommendations through experiments on skincare products (experience goods) across three commercial LLMs (GPT-4o-mini, Claude Sonnet, Gemini 3 Flash), with a robustness check on search goods. Key findings include: (1) a Conditional Monopoly in which well-known brands receive 100% of recommendations (IAI = 10.0) when all products have identical specifications, but this dominance vanishes with a competitor's rating advantage of less than +0.1 stars; (2) authority-style marketing language (including fabricated clinical claims) breaks the monopoly at a Bias Surplus Value equivalent to +0.17 rating points, with model-specific responses; and (3) a social dilemma in multi-brand GEO competition where universal adoption reduces individual payoffs from +0.802 to +0.007 in the payoff proxy, with non-participants receiving zero recommendations.

Significance. If the reported thresholds and collective-action effects hold, the work identifies concrete mechanisms by which brand reputation and optimization tactics shape LLM-mediated markets, with implications for competition policy and the study of generative engine optimization as both a security issue and a marketing practice. The multi-model design and category robustness check provide an empirical foundation that could be extended to quantify divergence across LLMs and product types.

major comments (1)
  1. [Experiments (including robustness check)] The headline claims concern behavior in 'LLM recommendation systems' broadly, yet rest on results from only three models and primarily the skincare category. The abstract notes a robustness check on search goods but provides no quantitative comparison of effect sizes, model divergence, or category-specific differences; without such data, the extension from the tested cases to the stated domain remains unsupported.
minor comments (1)
  1. [Abstract] The abstract summarizes results but omits essential methodological details (sample sizes, statistical tests, controls, exact prompt templates) that would allow readers to assess the experiments at a glance; these should be added even if elaborated in the main text.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the detailed and constructive report. We address the major comment on experimental scope and generalizability below, and commit to revisions that strengthen the manuscript without overstating current evidence.

read point-by-point responses
  1. Referee: [Experiments (including robustness check)] The headline claims concern behavior in 'LLM recommendation systems' broadly, yet rest on results from only three models and primarily the skincare category. The abstract notes a robustness check on search goods but provides no quantitative comparison of effect sizes, model divergence, or category-specific differences; without such data, the extension from the tested cases to the stated domain remains unsupported.

    Authors: We agree that the current manuscript does not provide sufficient quantitative detail on the robustness check to fully support broad claims. The three models tested are the leading commercial LLMs available during the study period, and skincare was selected as the primary category because it exemplifies experience goods where brand reputation is salient; the search-goods check was intended only as an initial robustness probe. In revision we will add a dedicated subsection (with tables) reporting direct comparisons of key metrics—IAI scores, monopoly rates, and Bias Surplus Values—between skincare and search goods, plus model-by-model effect sizes and divergence statistics. This will allow readers to assess the degree of consistency. We will also qualify the abstract and introduction to frame the results as initial evidence from leading models and two product types rather than a comprehensive survey of all LLM recommendation systems. revision: yes

Circularity Check

0 steps flagged

No circularity: empirical results from direct LLM experiments

full rationale

The paper reports controlled experiments measuring LLM recommendation frequencies for skincare products across three models. Central claims (Conditional Monopoly with IAI=10.0, bias surplus thresholds, social dilemma payoffs) are direct observational outputs from prompt-based tests, not derived via equations, fitted parameters, or self-citation chains. No load-bearing steps reduce to inputs by construction; the work is self-contained against external benchmarks.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The central claims rest on experimental measurements from three LLMs on skincare products; no new theoretical entities introduced. Review limited to abstract.

axioms (1)
  • domain assumption Skincare products represent experience goods where consumers rely on brand reputation.
    Used to justify choice of category in abstract.

pith-pipeline@v0.9.1-grok · 5776 in / 1313 out tokens · 46517 ms · 2026-06-27T01:37:37.733274+00:00 · methodology

0 comments
read the original abstract

Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel. We study brand dynamics in LLM recommendations using skincare products -- a category where consumers cannot easily judge quality before buying and must rely on brand reputation -- across three commercial LLMs (GPT-4o-mini, Claude Sonnet, Gemini 3 Flash), with a robustness check on search goods. In three experiments, we find: (1) a Conditional Monopoly where well-known brands get recommended 100% of the time (IAI = 10.0) when all products have the same specifications, but this dominance disappears with less than a +0.1-star rating advantage for a competitor; (2) authority-style marketing language, including fabricated clinical-evidence claims, breaks this monopoly at a Bias Surplus Value equal to +0.17 rating points, with each model responding differently; and (3) a social dilemma in multi-brand GEO competition: when all brands adopt the same optimization strategy, individual payoff falls from +0.802 to +0.007 in our payoff proxy, and non-participating brands receive zero recommendations in our tests. Our results suggest that generative engine optimization (GEO) should be studied not only as a security risk, but also as an emerging marketing practice that shapes market competition.

Figures

Figures reproduced from arXiv: 2606.17443 by Xi Chu, Yupeng Hou.

Figure 1
Figure 1. Figure 1: A real ChatGPT response (May 2026) to the [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Experiment 1 Summary—Conditional Monopoly. (a) IAI = 10.0: the real brand is recommended in 100% [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Experiment 2 Summary. (a) BR heatmap by bias [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Experiment 3 Summary. (a) U-shaped ISR trajectory. (b) Incumbent rank recovery. (c) Per-brand payoff [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

22 extracted references · 7 canonical work pages · 1 internal anchor

  1. [1]

    Amin Abolghasemi, Suzan Verberne, and Leif Azzopardi. 2024. Writing style matters: An examination of bias and fairness in information retrieval systems. ArXiv preprint arXiv:2411.13173

  2. [2]

    Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande. 2024. GEO : Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), pages 50--61

  3. [3]

    Bagga, Yuhang Wu, Pranjal Aggarwal, and Ameet Deshpande

    Puneet S. Bagga, Yuhang Wu, Pranjal Aggarwal, and Ameet Deshpande. 2025. E-GEO : A testbed for generative engine optimization in e-commerce. ArXiv preprint arXiv:2511.20867

  4. [4]

    Xu Chen, Zhehao Zhang, Yuqi Zhu, Mingyuan Tao, and Ji-Rong Wen. 2026. Auditing preferences for brands and cultures in LLMs . ArXiv preprint arXiv:2603.18300

  5. [5]

    Jessica Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, and Zhouhan Chen. 2024. Cognitive bias in decision-making with LLMs . In Findings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 12594--12613

  6. [6]

    Giorgos Filandrianos, Angeliki Dimitriou, Maria Lymperaiou, Konstantinos Thomas, and Giorgos Stamou. 2025. Bias beware: The impact of cognitive biases on LLM -driven product recommendations. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP)

  7. [7]

    Garrett Hardin. 1968. The tragedy of the commons. Science, 162(3859):1243--1248

  8. [8]

    Mahammed Kamruzzaman, Hieu Minh Nguyen, and Gene Louis Kim. 2024. `` Global is good, local is bad?'': Understanding brand bias in LLMs . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 12704--12721

  9. [9]

    Kevin Lane Keller. 1993. Conceptualizing, measuring, and managing customer-based brand equity. Journal of Marketing, 57(1):1--22

  10. [10]

    Aounon Kumar and Himabindu Lakkaraju. 2024. Manipulating large language models to increase product visibility. ArXiv preprint arXiv:2404.07981

  11. [11]

    Jan Malte Lichtenberg, Alexander Buchholz, and Pola Schw \"o bel. 2024. Large language models as recommender systems: A study of popularity bias. ArXiv preprint arXiv:2406.01285

  12. [12]

    Weiran Lin, Yizhong Wang, Lukas Bauer, Yun He, Minjoon Seo, and Daniel Khashabi. 2025. LLM whisperer: An inconspicuous attack to bias LLM responses. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI '25). ACM. Article 188375

  13. [13]

    Fatemeh Nazary, Yashar Deldjoo, and Tommaso Di Noia. 2025. Poison- RAG : Adversarial data poisoning attacks on retrieval-augmented generation in recommender systems. In Advances in Information Retrieval (ECIR 2025), pages 239--254. Springer

  14. [14]

    Phillip Nelson. 1970. Information and consumer behavior. Journal of Political Economy, 78(2):311--329

  15. [15]

    Fredrik Nestaas, Edoardo Debenedetti, and Florian Tram \`e r. 2025. Adversarial search engine optimization for large language models. In Proceedings of the International Conference on Learning Representations (ICLR)

  16. [16]

    OpenAI . 2026. Testing ads in ChatGPT . https://openai.com/index/testing-ads-in-chatgpt/. Published February 9, 2026. Accessed: 2026-05-03

  17. [17]

    Samuel Pfrommer, Yuze Bai, Tanmay Gautam, and Somayeh Sojoudi. 2024. Ranking manipulation for conversational search engines. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9520--9534

  18. [18]

    Talboy and Elizabeth Fuller

    Andr \'e s N. Talboy and Elizabeth Fuller. 2024. Challenging the status quo: Human bias in AI models? anchoring effects and mitigation strategies in large language models. Expert Systems with Applications, 255:124803

  19. [19]

    Lanling Xu, Junjie Zhang, Bingqian Yu, Jie Zhang, Hongzhi Li, Jingjing Li, and Xin Zhao. 2025. A survey on LLM -powered agents for recommender systems. In Findings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8015--8049

  20. [20]

    Jiaqi Xue, Mengxin Zheng, Yebowen Hu, Fei Liu, Xun Chen, and Qian Lou. 2024. BadRAG : Identifying vulnerabilities in retrieval-augmented generation of large language models. ArXiv preprint arXiv:2406.00083

  21. [21]

    Xikang Yang, Biyu Zhou, Xuehai Tang, Jizhong Han, and Songlin Hu. 2025. Exploiting synergistic cognitive biases to bypass safety in LLMs . ArXiv preprint arXiv:2507.22564

  22. [22]

    Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2025. PoisonedRAG : Knowledge corruption attacks to retrieval-augmented generation of large language models. In Proceedings of the 34th USENIX Security Symposium