REVIEW 3 major objections 5 minor 19 references
Legal puffery alone captures all the tool-selection bias that free-text registry copy can create; fabricated claims add nothing, and disclosure does not fix it.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 22:21 UTC pith:WWCIT6FI
load-bearing objection Strong multi-model measurement of description framing on tool selection; the first-best welfare claim for normalization is only as strong as the identical-tool premise the paper itself flags. the 3 major comments →
Agent-Facing Information Design in LLM Tool Registries
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Legal puffery alone produces the full optimization effect on agent tool selection (pooled legal uplift SBC ≈ +0.33); the illegal increment from fabricated claims is statistically zero. Superlatives are the single strongest feature (+0.35). Disclosure mechanisms fail model-dependently and cannot restore fair selection once models sit at ceiling. Registry-layer normalization of all descriptions to a functional L0 baseline eliminates the bias and achieves first-best welfare independently of the underlying model.
What carries the argument
Selection Bias Coefficient (SBC = P(select optimized) − 0.5) together with the framing-response function f that enters a Luce choice model; Propositions 1–3 show that high framing multipliers turn description optimization into a prisoners’ dilemma, that disclosure only partially attenuates κ, and that a registry normalization map ψ that rewrites every description to L0 restores the unique no-optimization equilibrium and first-best welfare.
Load-bearing premise
The welfare ranking of normalization over disclosure rests on treating the competing tools as functionally identical, so any money spent on framing is pure social waste; if richer copy actually tracks real capability, stripping it removes useful signal.
What would settle it
Re-run the legal-versus-full and structured-versus-optimized pairs on a live multi-tool registry where tools genuinely differ in measured capability (latency, coverage, accuracy) and check whether normalization still raises welfare or instead lowers selection of the truly better tool.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that LLM tool registries act as unregulated advertising markets: free-text provider descriptions drive agent tool selection, with no viewability, quality, or outcome infrastructure. Across 17,700+ trials (five LLMs, ten domains), it reports that commercial framing yields aggregate SBC ≈ +0.332, with four of five models at behavioral ceilings (P=1.0) in consumer domains; superlatives dominate single-feature ablation (SBC ≈ +0.35); legal puffery alone captures ≥100% of the optimization effect (E4, illegal increment ≈ −0.02); and disclosure (labels, ratings, system-prompt warnings) fails structurally for most models. A Luce-style strategic model (Propositions 1–3) frames description investment as a prisoner’s dilemma and claims that registry-layer normalization to L0 achieves first-best welfare model-independently. The constructive prescription separates selection-facing (structured, registry-controlled) from marketing-facing (provider-authored, post-selection) descriptions and introduces AAQS to separate capability from copywriting.
Significance. If the empirical pattern holds, the paper supplies the first systematic measurement of agent-facing commercial framing in tool registries and a concrete registry-design alternative to disclosure. Strengths include scale and design discipline: position-balanced ON/NN/OO cells, Wilson CIs, single-feature ablation (E2), legal-boundary spectrum (E4), multi-tool RSA (D), and partial ecological pairs (C); explicit calibration of framing multipliers κ into a standard choice model; and a falsifiable dual-description architecture plus AAQS. The finding that FTC-style deceptive-claim enforcement would miss the active mechanism (legal puffery) is policy-relevant. These contributions would matter for IR/agent-platform design even if the welfare ranking of normalization is scoped more carefully.
major comments (3)
- [§5 Setup / Prop. 3; Abstract; §6.3; §8] Prop. 3 and the Abstract/§1/§6.3 claim that registry normalization ψ:Q→{L0} “achieves first-best welfare model-independently” rest on the Setup assumption that tools are functionally identical (vi=v), so W=v−Σc(qi) and any framing cost is pure waste. §8 (“Identical-tool welfare assumption”) correctly flags the opposite case: when quality differs and description richness partially correlates with capability, normalization can destroy informative signaling. All primary experiments (A–E4) hold schemas and capability fixed by construction, so they never measure that tradeoff. The first-best ranking of normalization over disclosure/laissez-faire is therefore not established for heterogeneous registries the prescription is meant to govern. Either (i) qualify Prop. 3 and the abstract claim to the identical-tool case and state the ranking as conditional, or (ii) add analysis/bounds under positiv
- [§3.6 Experiment C; §4.5; Abstract contribution 5] Ecological support is preliminary and load-bearing for the claim that synthetic patterns transfer. Experiment C is n=360 across three domain pairs, with DCODE hybrid-synthetic and Claude DWEB position-confounded (Table 7). Two clean real pairs (Brave/DDG, Tomorrow.io/NWS) are useful but insufficient to underwrite “the synthetic pattern holds with naturally occurring descriptions” as a general result. Either expand C or clearly demote ecological claims to “preliminary evidence for two pairs” in the abstract and contribution list, and avoid treating C as confirmatory of the full design prescription.
- [§5 Eqs. (1)–(2); Prop. 1; Table 2; Fig. 3] The strategic model restricts providers to binary qi∈{0,q̄} matching L0 vs L3/L4 and uses IIA Luce choice (Eq. 1). Experiment B’s dose-response is non-monotone (L2 and L4 below L1/L3 for several models; Claude inverts at L4), and Experiment D shows winner-take-all only up to N=5. Prop. 1’s cost threshold and the “normalization is necessary” claim are calibrated to large κ, but the binary restriction and IIA are not stress-tested against continuous investment or variable-N dilution. A short robustness note—e.g., that Prop. 3 still holds under any f that is constant after ψ, while Prop. 1’s necessity claim is model- and κ-dependent—would keep the equilibrium story aligned with the non-monotone empirical response.
minor comments (5)
- [§4.1] Claude’s 74% last-position preference is acknowledged and partially corrected, but aggregate Claude SBC (+0.172) still appears in headline comparisons; always lead with position-controlled estimates when discussing model heterogeneity.
- [Table 1; Appendix C] Table 1 and the L0–L4 codebook (Table 9) are clear; consider releasing the full per-domain L0–L4 stimuli in the appendix or code release so E2/E4 can be fully reproduced without reverse-engineering.
- [§6.4 Eq. (3)] AAQS (Eq. 3) is a useful proposal, but Ci(q) is only sketched; a one-paragraph operational definition (e.g., binary structured-metadata match vs. LLM-judge) would make the score immediately usable rather than aspirational.
- [§8 Temperature] Temperature = 0 is stated as a limitation; a brief note on whether ceiling rates soften at T>0 (even on a subset) would help readers judge robustness of “behavioral ceiling” language.
- [Figure 1; §4.1] Figure 1’s “P(opt)=1.0 for 3/5 models” vs. text “four of five” in consumer domains is slightly inconsistent; align caption and body.
Circularity Check
No circular derivation: empirical SBCs are direct measurements; Propositions 1–3 are standard equilibrium consequences of a calibrated Luce model, not redefinitions of the data.
full rationale
The paper’s load-bearing claims are either (i) direct behavioral measurements (SBC, dose-response, E2 ablation, E4 legal-boundary, disclosure deltas) or (ii) equilibrium statements in a standard Luce choice model whose framing multiplier κ is calibrated from Experiment B and then used only to evaluate Nash thresholds. SBC is defined as P(select optimized)−0.5 and reported from trials; it is not fitted and then re-predicted. Proposition 3’s “first-best” result follows from the explicit identical-tool welfare definition W=v−Σc(qi) once normalization forces f(ψ(qi))=f(L0) and thus qi=0; that is a transparent model consequence under a stated assumption, not a self-definitional loop or a fitted quantity renamed as prediction. AAQS is introduced as a proposed score, not derived as a theorem. Citations (Luce 1959, Akerlof 1970, Cialdini, Edelman et al., prior MCP/GEO work) are external or non-overlapping; no uniqueness theorem or ansatz is imported from the present author to force the result. The identical-tool premise is a validity limitation (flagged in §8), not circularity. Score 0.
Axiom & Free-Parameter Ledger
free parameters (3)
- framing multiplier κ(Lk)=f(Lk)/f(L0) =
aggregate κ(L1)≈4.4, κ(L3)≈5.8; per-model table values
- illustrative revenue r and cost threshold c(q̄) =
threshold 0.159r at κ=5.8,N=5; r=$0.001/call illustrative
- framing level stimuli L0–L4 wording
axioms (5)
- domain assumption Agent selection follows Luce weights σi=f(qi)/Σf(qj) with IIA.
- ad hoc to paper When tools are functionally identical (vi=v), framing cost is pure social waste and W=v−Σc(qi).
- ad hoc to paper Providers choose binary framing qi∈{0,q̄} matching L0 vs L3/L4 experimental cells.
- domain assumption Temperature=0 deterministic tool choice represents agent selection behavior of interest.
- domain assumption Standard persuasion constructs (social proof, superlatives, outcome framing) are the right feature taxonomy for GEO-style tool copy.
invented entities (4)
-
SBC (selection bias coefficient) = P(select optimized)−0.5
independent evidence
-
Agent Attention Quality Score AAQSi(q)=P(select i|q)×Ci(q)
no independent evidence
-
Selection-facing vs marketing-facing dual description architecture
no independent evidence
-
Framing levels L0–L4 codebook
independent evidence
read the original abstract
LLM tool registries function as unregulated advertising platforms: providers write free-text descriptions that agents use for selection, yet no measurement infrastructure -- no viewability standard, quality score, or outcome audit -- exists to make this market accountable. We provide the first systematic framework, combining 17,700+ trials across five LLMs and ten domains with a constructive registry design prescription. Legal puffery alone (subjective superlatives, benefit framing) captures 100% of the optimization effect; fabricated claims add zero incremental bias -- rendering FTC enforcement of deceptive advertising rules ineffective against the active mechanism. Disclosure fails structurally: system-prompt warnings produce zero measurable effect for four of five models, and behavioral ceilings leave no headroom for label-based correction. Superlatives are the dominant single feature (SBC = +0.35). Registry-layer description normalization achieves first-best welfare model-independently. We propose separating selection-facing descriptions (structured, registry-controlled) from marketing-facing descriptions (provider-authored, shown post-selection), and introduce the Agent Attention Quality Score to distinguish capability from copywriting.
Figures
Reference graph
Works this paper leans on
-
[1]
P. Aggarwal et al. Generative engine optimization ( GEO ). arXiv:2311.09735, 2023
Pith/arXiv arXiv 2023
-
[2]
Agarwal, K
A. Agarwal, K. Hosanagar, and M. D. Smith. Do organic results help or hurt sponsored search performance? Information Systems Research, 22(4):837--859, 2011
2011
-
[3]
G. A. Akerlof. The market for ``lemons''. QJE, 84(3):488--500, 1970
1970
-
[4]
R. B. Cialdini. Influence: The Psychology of Persuasion. William Morrow, 1984
1984
-
[5]
A. Blankenstein et al. BiasBusters : Detecting and mitigating unintentional bias in LLM tool description quality. In ICLR 2026. arXiv:2510.00307
arXiv 2026
-
[6]
Edelman and D
B. Edelman and D. S. Gilchrist. Advertising disclosures: Measuring labeling alternatives in internet search engines. Info.\ Econ.\ & Policy, 24(1):75--89, 2012
2012
-
[7]
Edelman, M
B. Edelman, M. Ostrovsky, and M. Schwarz. Internet advertising and the generalized second-price auction. AER, 97(1):242--259, 2007
2007
-
[8]
F. Faghih et al. Bias beware: How cognitive biases manipulate LLM -based product recommendations. arXiv:2502.01349, 2025
arXiv 2025
-
[9]
Guides concerning the use of endorsements and testimonials in advertising (16 C.F.R
FTC . Guides concerning the use of endorsements and testimonials in advertising (16 C.F.R. Part 255), 2023
2023
-
[10]
M. Hasan et al. MCP tool descriptions are smelly. arXiv:2602.14878, 2026
Pith/arXiv arXiv 2026
-
[11]
G. N. Leech. English in Advertising: A Linguistic Study of Advertising in Great Britain. Longmans, 1966
1966
-
[12]
R. D. Luce. Individual Choice Behavior. Wiley, 1959
1959
-
[13]
B. Wang, Z. Liu, H. Yu, A. Yang, Y. Huang, J. Guo, H. Cheng, H. Li, and H. Wu. MCPGuard : Automatically detecting vulnerabilities in MCP servers. arXiv:2510.23673, 2025
arXiv 2025
-
[14]
G. Myers. Words in Ads. Edward Arnold, 1994
1994
-
[15]
P. R. Milgrom and R. J. Weber. A theory of auctions and competitive bidding. Econometrica, 50(5):1089--1122, 1982
1982
-
[16]
H. Wang, R. Zhang, J. Wang, M. Li, Y. Huang, D. Wang, and Q. Wang. ToolCommander : From allies to adversaries---manipulating LLM tool-calling through adversarial injection. arXiv:2412.10198, 2024
Pith/arXiv arXiv 2024
-
[17]
How online reviews influence sales
Spiegel Research Center . How online reviews influence sales. Northwestern University, 2017. https://spiegel.medill.northwestern.edu/online-reviews/. Accessed April 2026
2017
-
[18]
Tversky and D
A. Tversky and D. Kahneman. The framing of decisions and the psychology of choice. Science, 211(4481):453--458, 1981
1981
-
[19]
Y. Wang et al. MPMA : Manipulating LLM agents via MCP tool descriptions. arXiv:2505.11154, 2025
arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.