REVIEW 3 major objections 5 minor 22 references
Query language drives 26.5% of the variance in LLM brand answers, while brand identity drives only 1.5%.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 05:34 UTC pith:GI3MG2EA
load-bearing objection A transparent, well-scoped application of generalizability theory to LLM brand measurement, where the structural case against buying repeats is solid even though the headline variance split rests on a Gaussian fit to a 91.9%-neutral outcome. the 3 major comments →
Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On a fully crossed corpus of 12,933 responses about 20 brands in 8 languages from 3 LLMs, the paper fits a crossed random-effects model partitioning the variance of a single brand-sentiment response into brand, language, model, prompt, interactions, and residual. The central result: query language accounts for 26.5% of the variance of one response against 1.5% for brand identity, and after isolating pure resampling on the replicated subset, within-prompt resampling is 34.8% while the brand-in-context interaction is 29.6%. Object-by-facet, only brand-by-language is sizable (8.6%); brand-by-model and brand-by-prompt are zero. Feeding these components into the decision-study equation shows bran
What carries the argument
The crossed random-effects variance-components model (generalizability theory) with brand as object of measurement and language, model, and prompt as crossed random facets, plus a cell-level random intercept on the replicated subset so the residual becomes pure within-prompt resampling. The decision-study formula divides each variance component by the number of levels sampled for the facets it involves; because resampling is divided by the full query count while language is divided only by the language count, repeats have the fastest-decaying marginal value. The fitted components turn that structure into an explicit allocation rule.
Load-bearing premise
The whole variance partition assumes the sentiment scores, 91.9% of which are exactly zero, can be treated as Gaussian continuous measurements; if a zero-inflated or ordinal model gives a different decomposition, the 26.5%/1.5%/34.8% splits and the allocation rule could change.
What would settle it
Recompute M1–M3 on the same 12,933 responses with a binary or ordinal recommendation outcome (brand named, or its rank) instead of sentiment polarity. If brand identity's ICC rises well above 0.0146 and language's share falls below 26.5%, the paper's central quantitative claim is an artifact of the near-degenerate sentiment outcome. Alternatively, fit a zero-inflated model; if the language variance component drops materially, the Gaussian assumption is the culprit.
If this is right
- If true, any brand ranking built from a single LLM answer (or a few repeats in one language) is essentially noise; brand identity explains only 1.5% of single-response variance.
- Reliable brand measurement requires cross-language and cross-model breadth: 8 languages, 3 models, 1 paraphrase, 10 repeats reach Eρ² ≈ 0.13, whereas 20 repeats in one language/prompt reach only 0.02.
- The field-standard 'repeat the prompt five times and average' convention is the least efficient allocation of a query budget; paraphrase breadth (15 paraphrases, 1 repeat) beats repeat depth (5 paraphrases, 5 repeats) at lower cost.
- The reliability ceiling is low (~0.36) for this sentiment outcome; a less degenerate outcome such as a recommendation indicator is the natural next test.
- The structural ordering—repeats saturate first—holds for any positive variance components, so the conclusion generalizes beyond the fitted numbers.
Where Pith is reading between the lines
- The near-degenerate outcome (91.9% neutral) likely suppresses brand signal; a recommendation-indicator refit may raise ICCs but probably preserves the facet ordering, since the ordering is structural.
- The 26.5% language share is measured on sentiment polarity; on a binary 'is the brand named' outcome the language share could shrink or grow, and that is a direct testable extension from the same stored responses.
- A zero-inflated or ordinal model could change the variance split; if the split holds under that model, the conclusion is much stronger.
- The allocation rule has an immediate practical corollary for anyone tracking AI visibility: spend the next block of queries on a new language, not a sixth repeat.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper sets out to explain where non-determinism in LLM brand answers comes from, arguing that measured brand scores move for at least four separable reasons: within-prompt resampling, prompt paraphrase, model identity, and query language. It specifies a crossed random-effects (generalizability-theory) decomposition, fits three REML models (M1--M3) to a fully crossed corpus of 12,933 responses from 20 brands, 8 languages, and 3 models, and embeds the variance components in a decision-study allocation. The main empirical claims are that query language is the largest systematic facet (26.5% of single-response variance), brand identity is tiny (1.5%, ICC 0.0146), within-prompt resampling is 34.8% on a stability subset, brand-by-language is 8.6%, and that repeats beyond five are the least efficient use of a query budget. The paper is unusually transparent: it states in Sections 5.4 and 7 that cluster-bootstrap confidence intervals, a model-free within-cell cross-check, a recommendation-indicator refit, and the secondary-corpus fits are still to come.
Significance. If the estimates hold, the paper makes a useful practical and methodological contribution. It gives the field a variance-components vocabulary for a problem that is usually treated as 'resample five times,' and it provides a concrete, falsifiable allocation rule. The structural argument in Section 4.2---that the repeat facet's marginal value decays fastest because the resampling term is divided by the full query count---is sound and independent of the fitted numbers. The authors also ship runnable code and exact model specifications, which is a genuine strength. However, the headline numeric claims are point estimates from a Gaussian linear mixed model applied to an outcome that is 91.9% exactly neutral, with no confidence intervals and with a named model-free validation deferred. The significance is therefore conditional: the paper establishes a template and a structural prediction, but the specific 26.5%/1.5%/34.8% split and the 'repeats last' allocation are not yet secured by the evidence actually presented.
major comments (3)
- [Section 3.1, Section 4.1, Table 2, Limitation 2] The central claim that query language (26.5%) dwarfs brand identity (1.5%) is estimated by Gaussian REML on an outcome where 91.9% of responses are exactly neutral. A Gaussian likelihood on a point-mass-plus-tail mixture does not necessarily recover a variance partition of substantive brand signal; the language effect could be driven mainly by between-language differences in the probability of a neutral response, and the brand-by-language term by differences in zero rates rather than in sentiment among nonzero answers. The paper itself concedes this is a 'first-order approximation' and names a recommendation-indicator refit as future work. This is load-bearing because Tables 5 and 6, and all allocation conclusions, inherit the M3 components. I would need either a two-part (zero vs continuous) or ordinal GLMM refit, or at least a sensitivity analysis showing the facet ordering is stable u
- [Section 4.1, Section 5.2, Limitation 1] The claim that the M2 residual is pure within-prompt resampling (34.8%) is not yet validated. The paper states that the direct model-free mean within-cell variance cross-check is 'computed and deposited with the bootstrap intervals in the v2 finalization'---i.e., the validation result is not in the manuscript. Without that cross-check, the interpretation of the residual as pure resampling is an assumption of the model, not a demonstrated property. Additionally, resampling is identified only on Category D prompts, which are a subset of the 15 prompts; the paper acknowledges that transporting the split to the full corpus is an assumption. This matters because the 'repeats beyond five are least efficient' conclusion depends on the M2/M3 residual magnitude.
- [Section 4.2, Tables 5-6, Section 5.4] The decision-study frontier and allocation rule are deterministic functions of point estimates from a single REML fit. The cluster-bootstrap confidence intervals on all components, ICCs, and frontier coefficients are deferred to v2. With only 20 brands, those intervals are likely to be wide, and the paper itself says the point estimates should be read 'without interval guarantees until then.' Given that the abstract and discussion present the repeat-vs-language ordering as a settled conclusion, this is more than a presentation issue. I would require at least a sensitivity analysis over plausible component ranges, or the actual bootstrap intervals, before the allocation rule is reported as a finding.
minor comments (5)
- [Abstract and Section 5.2] The abstract states 'resampling is 34.8% of variance' without immediately noting that this is on the Category D stability subset, not the full corpus. The full-corpus residual is 69.3% and conflates resampling with interactions. Please make this conditional explicit in the abstract.
- [Table 2 and Table 4] Minor typographical issues: 'T otal' in Table 2 and the spacing in 'V ariance' in Table 1 should be fixed. More substantively, Table 4 reports brand-by-model and brand-by-prompt variances as exactly 0.000. These are almost certainly boundary estimates from the optimizer/powell fit and should be labeled as such, not reported as exact zeros.
- [Section 5.2] The sentence 'The direct model-free mean within-cell variance cross-check is computed and deposited with the bootstrap intervals in the v2 finalization' is ambiguous. If the value already exists in the repository, report it; if not, state clearly that it has not yet been computed.
- [Section 3.2-3.3] The two secondary corpora are described in the Data section but no decomposition is run on them. This is fine as a preview, but the section would be clearer if it explicitly said 'these corpora are not analyzed in this version' at the start of each subsection, rather than in a shared parenthetical.
- [References] Reference [17] is missing author names. Some references [18]-[22] are to the author's own preprints and industry reports; please mark which items are peer-reviewed, since Section 2 currently mixes them with peer-reviewed literature without distinction.
Circularity Check
No circular derivation chain; the decision-study allocation is arithmetic on fitted variance components, with only minor non-load-bearing self-citation.
full rationale
The central derivation is not circular. The variance components in M1–M3 are REML estimates obtained from the reported response-level corpus (Section 4.1, Tables 2–4), and the ICCs are their normalized shares. The decision-study frontier and the 'repeats past five are least efficient' allocation (Section 4.2, Tables 5–6) are obtained by substituting those fitted components into the standard generalizability-theory error-variance formula, σ2δ = σ2pL/nL + σ2pM/nM + σ2pP/nP + σ2cell/(nLnMnP) + σ2e/(nLnMnPnR); the paper itself labels Table 5 as 'Computed from the M3 components.' Thus the key numbers are arithmetic consequences of the fits, not inputs recycled as outputs. The 'structural prediction' that the repeat facet has the fastest-decaying marginal value is a property of that formula—resampling is divided by the full query count while language is divided only by the language count—so it is a mathematical consequence used to interpret the estimates, not a separate empirical claim used to set the constants. Self-citations [18]–[22] supply the dataset, the five-repeat convention, and adjacent analyses, but no load-bearing argument reduces to an unverified self-citation, and no uniqueness theorem is imported from the authors' prior work. The paper's own limitations—Gaussian REML on a 91.9%-neutral outcome (Limitation 2), missing cluster-bootstrap confidence intervals (Limitation 1), and the non-identifiability of brand×prompt against the cell term (Limitation 5)—are genuine correctness and misspecification risks that bear on the reliability of the point estimates, but they do not make the derivation equivalent to its inputs. The claimed decomposition and allocation therefore survive a circularity review, with the caveat that the 'structural prediction confirmed' language should be read as an arithmetic check rather than an independent confirmation.
Axiom & Free-Parameter Ledger
free parameters (1)
- REML variance-component estimates (M1-M3) =
sigma2_brand=0.000280; sigma2_lang=0.013518; cell=0.009522; resid=0.014873
axioms (5)
- domain assumption Random effects for brand, language, model, prompt and residual are independent, mean-zero Gaussian with constant variance.
- domain assumption The Category D stability subset identifies pure within-prompt resampling and the cell term, and this partition transports to the full corpus.
- domain assumption API collection missingness and retries are ignorable, with no systematic order or carryover effects.
- standard math The generalizability-theory error formulas for relative and absolute error variance are the correct decision-study specification.
- domain assumption Sentiment polarity is a meaningful continuous response-level outcome for the decomposition, despite noise in absolute calibration.
read the original abstract
Teams measuring whether large language models (LLMs) recommend a brand face a reproducibility problem: ask the same question twice and the answer moves. Practice resamples each prompt a few times (commonly five) and averages, treating within-prompt resampling as the source of the noise. But a measured brand score moves for at least four separable reasons: within-prompt resampling, prompt paraphrase, model identity, and query language. We specify a crossed random-effects (generalizability-theory) decomposition that partitions the total variance of a response-level brand outcome into these four sources, and embed the components in a decision-study allocation that returns how many repeats, paraphrases, models, and languages to buy for a target reliability. We apply it to a fully crossed corpus of 12,933 LLM responses on 20 Central and Eastern European brands, 8 languages, and 3 models (GPT-5.2 and Gemini 3 Flash in parametric mode, Perplexity in grounded retrieval), with a stability subset of 1,435 cells resampled about five times. The outcome is per-response multilingual sentiment polarity. Query language is the largest systematic facet (26.5% of the variance of one response) against 1.5% for brand identity (ICC 0.0146), so a single AI answer carries almost no brand-discriminating signal. Once a cell term isolates pure resampling, resampling is 34.8% of variance and the brand-in-context interaction 29.6%; brand-by-language is 8.6% (a bilingual penalty) while brand-by-model and brand-by-prompt are near zero. Per unit of query budget, adding languages and models reduces relative-error variance far more than adding repeats: a repeat past the fifth reduces it by only 0.0003. Brand-ranking reliability stays low, near 0.01 for a single answer and about 0.36 at the full crossed design, so reliability is bought by spreading across languages and models, not by repeating one prompt.
Reference graph
Works this paper leans on
-
[1]
Atil, B., et al. (2025). Non-Determinism of “Deterministic” LLM System Set- tings in Hosted Environments. InProceedings of the 5th Workshop on Evaluation and Comparison of NLP Systems (Eval4NLP), EMNLP 2025. ACL Anthology 2025.eval4nlp-1.12 (arXiv:2408.04667)
Pith/arXiv arXiv 2025
-
[2]
Barbieri, F., Espinosa Anke, L., and Camacho-Collados, J. (2022). XLM-T: Multilingual language models in Twitter for sentiment analysis and beyond. In Proceedings of LREC 2022, 258 to 266
2022
-
[3]
Bates, D., M¨ achler, M., Bolker, B., and Walker, S. (2015). Fitting linear mixed- effects models using lme4.Journal of Statistical Software, 67(1), 1 to 48
2015
-
[4]
Brennan, R. L. (2001).Generalizability Theory. Springer
2001
-
[5]
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzm´ an, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. (2020). Unsupervised cross-lingual representation learning at scale. InProceedings of ACL 2020, 8440 to 8451
2020
-
[6]
J., Gleser, G
Cronbach, L. J., Gleser, G. C., Nanda, H., and Rajaratnam, N. (1972).The Dependability of Behavioral Measurements: Theory of Generalizability for Scores and Profiles. Wiley
1972
-
[7]
Efron, B., and Tibshirani, R. J. (1993).An Introduction to the Bootstrap. Chapman and Hall
1993
-
[8]
How many prompts and responses you need for a stable AI position estimate
Graphite (2026). How many prompts and responses you need for a stable AI position estimate. Industry report, May 2026 (200 prompts by 400 responses). Referenced via Search Engine Land
2026
-
[9]
Indig, K. (2026). Measuring AI search: run-to-run citation stability. Referenced via Search Engine Land, June 2026
2026
-
[10]
Nakagawa, S., and Schielzeth, H. (2013). A general and simple method for obtain- ing R2 from generalized linear mixed-effects models.Methods in Ecology and Evolution, 4(2), 133 to 142
2013
-
[11]
Ouyang, S., Zhang, J. M., Harman, M., and Wang, M. (2025). An empirical study of the non-determinism of ChatGPT in code generation.ACM Transactions on Software Engineering and Methodology(arXiv:2308.02828)
Pith/arXiv arXiv 2025
-
[12]
Fan-out and citation measurement reports
Peec AI (2026). Fan-out and citation measurement reports. Industry reports, February to May 2026
2026
-
[13]
Citation concentration in AI answers
Profound (2025). Citation concentration in AI answers. Industry reports, June and November 2025. 17
2025
-
[14]
R., Casella, G., and McCulloch, C
Searle, S. R., Casella, G., and McCulloch, C. E. (1992).Variance Components. Wiley
1992
-
[15]
Seabold, S., and Perktold, J. (2010). statsmodels: Econometric and statistical modeling with Python. InProceedings of the 9th Python in Science Conference
2010
-
[16]
J., and Webb, N
Shavelson, R. J., and Webb, N. M. (1991).Generalizability Theory: A Primer. Sage
1991
-
[17]
Understanding and mitigating numerical sources of nondeterminism in LLM inference (2025). arXiv:2506.09501
arXiv 2025
-
[18]
Zatuchin, D. (2026a). The category-ownership map: measuring which brands own which recommendation categories in AI. Dataset DOI 10.5281/zenodo.20788142. Preprint; under review at the Journal of Marketing Analytics
-
[19]
Zatuchin, D. (2026b). The dice-roll method: power and generalizability for measuring LLM brand recommendations. Preprint; under peer review
-
[20]
Zatuchin, D. (2026c). Cross-language AI brand reputation across European lan- guages, including the Central and Eastern European arm used here. Dataset DOI 10.5281/zenodo.20794390. Preprint
-
[21]
Zatuchin, D. (2026d). Six corporates, three visibility patterns: a dice-roll audit of large-cap AI visibility. Rankfor.AI research article, open.rankfor.ai, 2026
2026
-
[22]
Zatuchin, D. (2026e). V-ZUG and the incumbent shadow: unbranded-query visibility in premium appliances. Rankfor.AI research article, open.rankfor.ai, 2026. 18
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.