Pith. sign in

REVIEW 3 major objections 5 minor 22 references

Query language drives 26.5% of the variance in LLM brand answers, while brand identity drives only 1.5%.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 05:34 UTC pith:GI3MG2EA

load-bearing objection A transparent, well-scoped application of generalizability theory to LLM brand measurement, where the structural case against buying repeats is solid even though the headline variance split rests on a Gaussian fit to a 91.9%-neutral outcome. the 3 major comments →

arxiv 2607.13304 v1 pith:GI3MG2EA submitted 2026-07-14 cs.IR cs.CL

Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers

classification cs.IR cs.CL
keywords generalizability theoryvariance componentsintraclass correlationlarge language modelsmeasurement reproducibilitysampling designbrand measurementquery language
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish where the run-to-run movement in LLM brand answers comes from, separating four sources: resampling the same prompt, paraphrasing the prompt, switching models, and asking in a different language. On its corpus, query language is by far the largest systematic source (26.5% of single-response variance) while brand identity is only 1.5% (ICC 0.0146), so one answer from one prompt in one language carries almost no information about where a brand ranks. The paper then uses the variance components to derive a spending rule: for a fixed query budget, adding languages and models reduces ranking error far more than adding more repeats, and repeats beyond five are the least efficient use of a query. If true, the common practice of averaging five repetitions of the same prompt is the wrong place to spend measurement effort.

Core claim

On a fully crossed corpus of 12,933 responses about 20 brands in 8 languages from 3 LLMs, the paper fits a crossed random-effects model partitioning the variance of a single brand-sentiment response into brand, language, model, prompt, interactions, and residual. The central result: query language accounts for 26.5% of the variance of one response against 1.5% for brand identity, and after isolating pure resampling on the replicated subset, within-prompt resampling is 34.8% while the brand-in-context interaction is 29.6%. Object-by-facet, only brand-by-language is sizable (8.6%); brand-by-model and brand-by-prompt are zero. Feeding these components into the decision-study equation shows bran

What carries the argument

The crossed random-effects variance-components model (generalizability theory) with brand as object of measurement and language, model, and prompt as crossed random facets, plus a cell-level random intercept on the replicated subset so the residual becomes pure within-prompt resampling. The decision-study formula divides each variance component by the number of levels sampled for the facets it involves; because resampling is divided by the full query count while language is divided only by the language count, repeats have the fastest-decaying marginal value. The fitted components turn that structure into an explicit allocation rule.

Load-bearing premise

The whole variance partition assumes the sentiment scores, 91.9% of which are exactly zero, can be treated as Gaussian continuous measurements; if a zero-inflated or ordinal model gives a different decomposition, the 26.5%/1.5%/34.8% splits and the allocation rule could change.

What would settle it

Recompute M1–M3 on the same 12,933 responses with a binary or ordinal recommendation outcome (brand named, or its rank) instead of sentiment polarity. If brand identity's ICC rises well above 0.0146 and language's share falls below 26.5%, the paper's central quantitative claim is an artifact of the near-degenerate sentiment outcome. Alternatively, fit a zero-inflated model; if the language variance component drops materially, the Gaussian assumption is the culprit.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If true, any brand ranking built from a single LLM answer (or a few repeats in one language) is essentially noise; brand identity explains only 1.5% of single-response variance.
  • Reliable brand measurement requires cross-language and cross-model breadth: 8 languages, 3 models, 1 paraphrase, 10 repeats reach Eρ² ≈ 0.13, whereas 20 repeats in one language/prompt reach only 0.02.
  • The field-standard 'repeat the prompt five times and average' convention is the least efficient allocation of a query budget; paraphrase breadth (15 paraphrases, 1 repeat) beats repeat depth (5 paraphrases, 5 repeats) at lower cost.
  • The reliability ceiling is low (~0.36) for this sentiment outcome; a less degenerate outcome such as a recommendation indicator is the natural next test.
  • The structural ordering—repeats saturate first—holds for any positive variance components, so the conclusion generalizes beyond the fitted numbers.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The near-degenerate outcome (91.9% neutral) likely suppresses brand signal; a recommendation-indicator refit may raise ICCs but probably preserves the facet ordering, since the ordering is structural.
  • The 26.5% language share is measured on sentiment polarity; on a binary 'is the brand named' outcome the language share could shrink or grow, and that is a direct testable extension from the same stored responses.
  • A zero-inflated or ordinal model could change the variance split; if the split holds under that model, the conclusion is much stronger.
  • The allocation rule has an immediate practical corollary for anyone tracking AI visibility: spend the next block of queries on a new language, not a sixth repeat.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper sets out to explain where non-determinism in LLM brand answers comes from, arguing that measured brand scores move for at least four separable reasons: within-prompt resampling, prompt paraphrase, model identity, and query language. It specifies a crossed random-effects (generalizability-theory) decomposition, fits three REML models (M1--M3) to a fully crossed corpus of 12,933 responses from 20 brands, 8 languages, and 3 models, and embeds the variance components in a decision-study allocation. The main empirical claims are that query language is the largest systematic facet (26.5% of single-response variance), brand identity is tiny (1.5%, ICC 0.0146), within-prompt resampling is 34.8% on a stability subset, brand-by-language is 8.6%, and that repeats beyond five are the least efficient use of a query budget. The paper is unusually transparent: it states in Sections 5.4 and 7 that cluster-bootstrap confidence intervals, a model-free within-cell cross-check, a recommendation-indicator refit, and the secondary-corpus fits are still to come.

Significance. If the estimates hold, the paper makes a useful practical and methodological contribution. It gives the field a variance-components vocabulary for a problem that is usually treated as 'resample five times,' and it provides a concrete, falsifiable allocation rule. The structural argument in Section 4.2---that the repeat facet's marginal value decays fastest because the resampling term is divided by the full query count---is sound and independent of the fitted numbers. The authors also ship runnable code and exact model specifications, which is a genuine strength. However, the headline numeric claims are point estimates from a Gaussian linear mixed model applied to an outcome that is 91.9% exactly neutral, with no confidence intervals and with a named model-free validation deferred. The significance is therefore conditional: the paper establishes a template and a structural prediction, but the specific 26.5%/1.5%/34.8% split and the 'repeats last' allocation are not yet secured by the evidence actually presented.

major comments (3)
  1. [Section 3.1, Section 4.1, Table 2, Limitation 2] The central claim that query language (26.5%) dwarfs brand identity (1.5%) is estimated by Gaussian REML on an outcome where 91.9% of responses are exactly neutral. A Gaussian likelihood on a point-mass-plus-tail mixture does not necessarily recover a variance partition of substantive brand signal; the language effect could be driven mainly by between-language differences in the probability of a neutral response, and the brand-by-language term by differences in zero rates rather than in sentiment among nonzero answers. The paper itself concedes this is a 'first-order approximation' and names a recommendation-indicator refit as future work. This is load-bearing because Tables 5 and 6, and all allocation conclusions, inherit the M3 components. I would need either a two-part (zero vs continuous) or ordinal GLMM refit, or at least a sensitivity analysis showing the facet ordering is stable u
  2. [Section 4.1, Section 5.2, Limitation 1] The claim that the M2 residual is pure within-prompt resampling (34.8%) is not yet validated. The paper states that the direct model-free mean within-cell variance cross-check is 'computed and deposited with the bootstrap intervals in the v2 finalization'---i.e., the validation result is not in the manuscript. Without that cross-check, the interpretation of the residual as pure resampling is an assumption of the model, not a demonstrated property. Additionally, resampling is identified only on Category D prompts, which are a subset of the 15 prompts; the paper acknowledges that transporting the split to the full corpus is an assumption. This matters because the 'repeats beyond five are least efficient' conclusion depends on the M2/M3 residual magnitude.
  3. [Section 4.2, Tables 5-6, Section 5.4] The decision-study frontier and allocation rule are deterministic functions of point estimates from a single REML fit. The cluster-bootstrap confidence intervals on all components, ICCs, and frontier coefficients are deferred to v2. With only 20 brands, those intervals are likely to be wide, and the paper itself says the point estimates should be read 'without interval guarantees until then.' Given that the abstract and discussion present the repeat-vs-language ordering as a settled conclusion, this is more than a presentation issue. I would require at least a sensitivity analysis over plausible component ranges, or the actual bootstrap intervals, before the allocation rule is reported as a finding.
minor comments (5)
  1. [Abstract and Section 5.2] The abstract states 'resampling is 34.8% of variance' without immediately noting that this is on the Category D stability subset, not the full corpus. The full-corpus residual is 69.3% and conflates resampling with interactions. Please make this conditional explicit in the abstract.
  2. [Table 2 and Table 4] Minor typographical issues: 'T otal' in Table 2 and the spacing in 'V ariance' in Table 1 should be fixed. More substantively, Table 4 reports brand-by-model and brand-by-prompt variances as exactly 0.000. These are almost certainly boundary estimates from the optimizer/powell fit and should be labeled as such, not reported as exact zeros.
  3. [Section 5.2] The sentence 'The direct model-free mean within-cell variance cross-check is computed and deposited with the bootstrap intervals in the v2 finalization' is ambiguous. If the value already exists in the repository, report it; if not, state clearly that it has not yet been computed.
  4. [Section 3.2-3.3] The two secondary corpora are described in the Data section but no decomposition is run on them. This is fine as a preview, but the section would be clearer if it explicitly said 'these corpora are not analyzed in this version' at the start of each subsection, rather than in a shared parenthetical.
  5. [References] Reference [17] is missing author names. Some references [18]-[22] are to the author's own preprints and industry reports; please mark which items are peer-reviewed, since Section 2 currently mixes them with peer-reviewed literature without distinction.

Circularity Check

0 steps flagged

No circular derivation chain; the decision-study allocation is arithmetic on fitted variance components, with only minor non-load-bearing self-citation.

full rationale

The central derivation is not circular. The variance components in M1–M3 are REML estimates obtained from the reported response-level corpus (Section 4.1, Tables 2–4), and the ICCs are their normalized shares. The decision-study frontier and the 'repeats past five are least efficient' allocation (Section 4.2, Tables 5–6) are obtained by substituting those fitted components into the standard generalizability-theory error-variance formula, σ2δ = σ2pL/nL + σ2pM/nM + σ2pP/nP + σ2cell/(nLnMnP) + σ2e/(nLnMnPnR); the paper itself labels Table 5 as 'Computed from the M3 components.' Thus the key numbers are arithmetic consequences of the fits, not inputs recycled as outputs. The 'structural prediction' that the repeat facet has the fastest-decaying marginal value is a property of that formula—resampling is divided by the full query count while language is divided only by the language count—so it is a mathematical consequence used to interpret the estimates, not a separate empirical claim used to set the constants. Self-citations [18]–[22] supply the dataset, the five-repeat convention, and adjacent analyses, but no load-bearing argument reduces to an unverified self-citation, and no uniqueness theorem is imported from the authors' prior work. The paper's own limitations—Gaussian REML on a 91.9%-neutral outcome (Limitation 2), missing cluster-bootstrap confidence intervals (Limitation 1), and the non-identifiability of brand×prompt against the cell term (Limitation 5)—are genuine correctness and misspecification risks that bear on the reliability of the point estimates, but they do not make the derivation equivalent to its inputs. The claimed decomposition and allocation therefore survive a circularity review, with the caveat that the 'structural prediction confirmed' language should be read as an arithmetic check rather than an independent confirmation.

Axiom & Free-Parameter Ledger

1 free parameters · 5 axioms · 0 invented entities

The central estimates ride on standard g-theory plus REML assumptions and on three data-domain assumptions the paper itself flags: Gaussian model on a near-degenerate outcome, stability-subset transport, and ignorable missingness. No invented entities are posited. The fitted variance components are empirical estimates, not ad hoc constants.

free parameters (1)
  • REML variance-component estimates (M1-M3) = sigma2_brand=0.000280; sigma2_lang=0.013518; cell=0.009522; resid=0.014873
    The ICCs, frontier, and allocation rule are arithmetic functions of these fitted components. They are estimated on this corpus, not derived from first principles or independently predicted.
axioms (5)
  • domain assumption Random effects for brand, language, model, prompt and residual are independent, mean-zero Gaussian with constant variance.
    Used in all REML fits (Section 4.1); questionable because 91.9% of outcomes are exactly neutral (Section 3.1, Limitation 2).
  • domain assumption The Category D stability subset identifies pure within-prompt resampling and the cell term, and this partition transports to the full corpus.
    M2 and M3 are fit on the Category D subset only; single-shot cells have no replication, so M1's residual split rests on representativeness (Section 5.2).
  • domain assumption API collection missingness and retries are ignorable, with no systematic order or carryover effects.
    The design issued 14,400 API calls versus 12,960 target responses; the 27 missing responses are treated as ignorable (Section 3.1).
  • standard math The generalizability-theory error formulas for relative and absolute error variance are the correct decision-study specification.
    Used in Section 4.2 to derive the allocation rule; this is standard g-theory machinery.
  • domain assumption Sentiment polarity is a meaningful continuous response-level outcome for the decomposition, despite noise in absolute calibration.
    The paper states the outcome is zero-inflated and noisy in absolute terms, and that the decomposition concerns relative variation only (Section 3.1, Limitation 2).

pith-pipeline@v1.3.0-alltime-deepseek · 13236 in / 20128 out tokens · 202473 ms · 2026-08-02T05:34:56.187807+00:00 · methodology

0 comments
read the original abstract

Teams measuring whether large language models (LLMs) recommend a brand face a reproducibility problem: ask the same question twice and the answer moves. Practice resamples each prompt a few times (commonly five) and averages, treating within-prompt resampling as the source of the noise. But a measured brand score moves for at least four separable reasons: within-prompt resampling, prompt paraphrase, model identity, and query language. We specify a crossed random-effects (generalizability-theory) decomposition that partitions the total variance of a response-level brand outcome into these four sources, and embed the components in a decision-study allocation that returns how many repeats, paraphrases, models, and languages to buy for a target reliability. We apply it to a fully crossed corpus of 12,933 LLM responses on 20 Central and Eastern European brands, 8 languages, and 3 models (GPT-5.2 and Gemini 3 Flash in parametric mode, Perplexity in grounded retrieval), with a stability subset of 1,435 cells resampled about five times. The outcome is per-response multilingual sentiment polarity. Query language is the largest systematic facet (26.5% of the variance of one response) against 1.5% for brand identity (ICC 0.0146), so a single AI answer carries almost no brand-discriminating signal. Once a cell term isolates pure resampling, resampling is 34.8% of variance and the brand-in-context interaction 29.6%; brand-by-language is 8.6% (a bilingual penalty) while brand-by-model and brand-by-prompt are near zero. Per unit of query budget, adding languages and models reduces relative-error variance far more than adding repeats: a repeat past the fifth reduces it by only 0.0003. Brand-ranking reliability stays low, near 0.01 for a single answer and about 0.36 at the full crossed design, so reliability is bought by spreading across languages and models, not by repeating one prompt.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

22 extracted references · 2 canonical work pages

  1. [1]

    Deterministic

    Atil, B., et al. (2025). Non-Determinism of “Deterministic” LLM System Set- tings in Hosted Environments. InProceedings of the 5th Workshop on Evaluation and Comparison of NLP Systems (Eval4NLP), EMNLP 2025. ACL Anthology 2025.eval4nlp-1.12 (arXiv:2408.04667)

  2. [2]

    Barbieri, F., Espinosa Anke, L., and Camacho-Collados, J. (2022). XLM-T: Multilingual language models in Twitter for sentiment analysis and beyond. In Proceedings of LREC 2022, 258 to 266

  3. [3]

    Bates, D., M¨ achler, M., Bolker, B., and Walker, S. (2015). Fitting linear mixed- effects models using lme4.Journal of Statistical Software, 67(1), 1 to 48

  4. [4]

    Brennan, R. L. (2001).Generalizability Theory. Springer

  5. [5]

    Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzm´ an, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. (2020). Unsupervised cross-lingual representation learning at scale. InProceedings of ACL 2020, 8440 to 8451

  6. [6]

    J., Gleser, G

    Cronbach, L. J., Gleser, G. C., Nanda, H., and Rajaratnam, N. (1972).The Dependability of Behavioral Measurements: Theory of Generalizability for Scores and Profiles. Wiley

  7. [7]

    Efron, B., and Tibshirani, R. J. (1993).An Introduction to the Bootstrap. Chapman and Hall

  8. [8]

    How many prompts and responses you need for a stable AI position estimate

    Graphite (2026). How many prompts and responses you need for a stable AI position estimate. Industry report, May 2026 (200 prompts by 400 responses). Referenced via Search Engine Land

  9. [9]

    Indig, K. (2026). Measuring AI search: run-to-run citation stability. Referenced via Search Engine Land, June 2026

  10. [10]

    Nakagawa, S., and Schielzeth, H. (2013). A general and simple method for obtain- ing R2 from generalized linear mixed-effects models.Methods in Ecology and Evolution, 4(2), 133 to 142

  11. [11]

    M., Harman, M., and Wang, M

    Ouyang, S., Zhang, J. M., Harman, M., and Wang, M. (2025). An empirical study of the non-determinism of ChatGPT in code generation.ACM Transactions on Software Engineering and Methodology(arXiv:2308.02828)

  12. [12]

    Fan-out and citation measurement reports

    Peec AI (2026). Fan-out and citation measurement reports. Industry reports, February to May 2026

  13. [13]

    Citation concentration in AI answers

    Profound (2025). Citation concentration in AI answers. Industry reports, June and November 2025. 17

  14. [14]

    R., Casella, G., and McCulloch, C

    Searle, S. R., Casella, G., and McCulloch, C. E. (1992).Variance Components. Wiley

  15. [15]

    Seabold, S., and Perktold, J. (2010). statsmodels: Econometric and statistical modeling with Python. InProceedings of the 9th Python in Science Conference

  16. [16]

    J., and Webb, N

    Shavelson, R. J., and Webb, N. M. (1991).Generalizability Theory: A Primer. Sage

  17. [17]

    arXiv:2506.09501

    Understanding and mitigating numerical sources of nondeterminism in LLM inference (2025). arXiv:2506.09501

  18. [18]

    Zatuchin, D. (2026a). The category-ownership map: measuring which brands own which recommendation categories in AI. Dataset DOI 10.5281/zenodo.20788142. Preprint; under review at the Journal of Marketing Analytics

  19. [19]

    Zatuchin, D. (2026b). The dice-roll method: power and generalizability for measuring LLM brand recommendations. Preprint; under peer review

  20. [20]

    Zatuchin, D. (2026c). Cross-language AI brand reputation across European lan- guages, including the Central and Eastern European arm used here. Dataset DOI 10.5281/zenodo.20794390. Preprint

  21. [21]

    Zatuchin, D. (2026d). Six corporates, three visibility patterns: a dice-roll audit of large-cap AI visibility. Rankfor.AI research article, open.rankfor.ai, 2026

  22. [22]

    Zatuchin, D. (2026e). V-ZUG and the incumbent shadow: unbranded-query visibility in premium appliances. Rankfor.AI research article, open.rankfor.ai, 2026. 18