{"id":"f99a1d84-4053-423f-805a-8ca7da5b993a","arxiv_id":"2607.13304","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"For AI brand answers, query language is the biggest source of answer noise (26.5%) while brand identity is tiny (1.5%), so repeating the same prompt is the least efficient way to spend a query budget.","lead":"This paper measures why AI answers about brands change between repetitions, splitting the answer-to-answer noise into four sources: re-running the same prompt, rephrasing, changing model, and changing language. It finds query language dominates and brand identity contributes almost nothing, so diversifying languages and models buys more reliability than repeating prompts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gaussian REML on a 91.9%-neutral outcome may manufacture the language-vs-brand split and the 'repeats last' allocation; the two-part/ordinal refit is future work, so the central claim is unsecured.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the Gaussian linear mixed model is fit to an outcome where 91.9% of responses are exactly neutral, and the resulting variance components drive the entire decision-study conclusion. I agree, and would sharpen the point: the problem is not merely that the residual is non-normal; it is that the dominant systematic facet (language) may be an artifact of the scorer's neutral mass. If the true data-generating process is a two-part mixture or ordinal, the reported 26.5% versus 1.5% split and the 'repeats beyond five are least efficient' rule could change materially. The paper is transparent about this being a first-order decomposition and defers the recommendation-indicator refit, but the abstract and discussion present the split as a finding. The structural algebra in Section 4.2 is correct given the components; the weak point is the mapping from data to components. A two-part/ordinal refit on the same response-level table is a concrete, low-cost check that would settle whether the concern lands. Since the paper itself identifies this as remaining work, the verdict should remain CONDITIONAL, not be upgraded or rejected.","tokens_in":13603,"tokens_out":6479,"duration_ms":87901,"concrete_test":"Fit a two-part refit on the Category D response-level table (available from author per Section 8): a logistic mixed model for the zero/non-zero split and a Gaussian mixed model for non-zero values, with the same crossed random effects as M3, plus a cumulative-link version on binned scores. Compare (a) the implied ICCs for language, brand, and brand-by-language, and (b) the re-computed Table 5 marginal reductions and Table 6 frontier. If the language share drops by more than about 5 percentage points, brand ICC rises above roughly 0.05, or the 'languages/models before repeats' ordering reverses, the Gaussian components do not support the abstract claim as stated. If the ordering and shares are stable, the concern is refuted and the CONDITIONAL verdict can be upgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline numbers (26.5% language, 1.5% brand, 8.6% brand-by-language, and the Table 5/6 allocation in which repeats past five are least efficient) all come from Gaussian REML fits on an outcome where 91.9% of responses are exactly 0 (Section 3.1). A Gaussian likelihood on a point-mass-plus-tail distribution does not estimate a variance partition of the substantive brand signal; it fits first and second moments of a mixture. Language effects could be carried mainly by between-language differences in the probability of a neutral response, and brand-by-language could be a difference in zero rates rather than in sentiment among non-neutral answers. Because the decision-study formulas in Section 4.2 feed these components directly into the 'repeats are the least efficient spend' conclusion, the central practical claim is conditional on the Gaussian approximation. The paper concedes this in Limitation 2 ('first-order approximation', 'ICCs are outcome-conditional') and Section 5.4 names the recommendation-indicator refit as future work, but the abstract and discussion state the 26.5%/1.5% split and the repeat allocation without that caveat. This is not an internal inconsistency; it is an unresolved misspecification risk at the point where the strongest claim is anchored. The promised model-free within-cell cross-check (Section 5.2) would validate the resampling variance only, not the language/brand split.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper sets out to explain where non-determinism in LLM brand answers comes from, arguing that measured brand scores move for at least four separable reasons: within-prompt resampling, prompt paraphrase, model identity, and query language. It specifies a crossed random-effects (generalizability-theory) decomposition, fits three REML models (M1--M3) to a fully crossed corpus of 12,933 responses from 20 brands, 8 languages, and 3 models, and embeds the variance components in a decision-study allocation. The main empirical claims are that query language is the largest systematic facet (26.5% of single-response variance), brand identity is tiny (1.5%, ICC 0.0146), within-prompt resampling is 34.8% on a stability subset, brand-by-language is 8.6%, and that repeats beyond five are the least efficient use of a query budget. The paper is unusually transparent: it states in Sections 5.4 and 7 that cluster-bootstrap confidence intervals, a model-free within-cell cross-check, a recommendation-indicator refit, and the secondary-corpus fits are still to come.","tokens_in":13888,"tokens_out":6463,"duration_ms":93161,"significance":"If the estimates hold, the paper makes a useful practical and methodological contribution. It gives the field a variance-components vocabulary for a problem that is usually treated as 'resample five times,' and it provides a concrete, falsifiable allocation rule. The structural argument in Section 4.2---that the repeat facet's marginal value decays fastest because the resampling term is divided by the full query count---is sound and independent of the fitted numbers. The authors also ship runnable code and exact model specifications, which is a genuine strength. However, the headline numeric claims are point estimates from a Gaussian linear mixed model applied to an outcome that is 91.9% exactly neutral, with no confidence intervals and with a named model-free validation deferred. The significance is therefore conditional: the paper establishes a template and a structural prediction, but the specific 26.5%/1.5%/34.8% split and the 'repeats last' allocation are not yet secured by the evidence actually presented.","major_comments":[{"comment":"The central claim that query language (26.5%) dwarfs brand identity (1.5%) is estimated by Gaussian REML on an outcome where 91.9% of responses are exactly neutral. A Gaussian likelihood on a point-mass-plus-tail mixture does not necessarily recover a variance partition of substantive brand signal; the language effect could be driven mainly by between-language differences in the probability of a neutral response, and the brand-by-language term by differences in zero rates rather than in sentiment among nonzero answers. The paper itself concedes this is a 'first-order approximation' and names a recommendation-indicator refit as future work. This is load-bearing because Tables 5 and 6, and all allocation conclusions, inherit the M3 components. I would need either a two-part (zero vs continuous) or ordinal GLMM refit, or at least a sensitivity analysis showing the facet ordering is stable u","section":"Section 3.1, Section 4.1, Table 2, Limitation 2"},{"comment":"The claim that the M2 residual is pure within-prompt resampling (34.8%) is not yet validated. The paper states that the direct model-free mean within-cell variance cross-check is 'computed and deposited with the bootstrap intervals in the v2 finalization'---i.e., the validation result is not in the manuscript. Without that cross-check, the interpretation of the residual as pure resampling is an assumption of the model, not a demonstrated property. Additionally, resampling is identified only on Category D prompts, which are a subset of the 15 prompts; the paper acknowledges that transporting the split to the full corpus is an assumption. This matters because the 'repeats beyond five are least efficient' conclusion depends on the M2/M3 residual magnitude.","section":"Section 4.1, Section 5.2, Limitation 1"},{"comment":"The decision-study frontier and allocation rule are deterministic functions of point estimates from a single REML fit. The cluster-bootstrap confidence intervals on all components, ICCs, and frontier coefficients are deferred to v2. With only 20 brands, those intervals are likely to be wide, and the paper itself says the point estimates should be read 'without interval guarantees until then.' Given that the abstract and discussion present the repeat-vs-language ordering as a settled conclusion, this is more than a presentation issue. I would require at least a sensitivity analysis over plausible component ranges, or the actual bootstrap intervals, before the allocation rule is reported as a finding.","section":"Section 4.2, Tables 5-6, Section 5.4"}],"minor_comments":[{"comment":"The abstract states 'resampling is 34.8% of variance' without immediately noting that this is on the Category D stability subset, not the full corpus. The full-corpus residual is 69.3% and conflates resampling with interactions. Please make this conditional explicit in the abstract.","section":"Abstract and Section 5.2"},{"comment":"Minor typographical issues: 'T otal' in Table 2 and the spacing in 'V ariance' in Table 1 should be fixed. More substantively, Table 4 reports brand-by-model and brand-by-prompt variances as exactly 0.000. These are almost certainly boundary estimates from the optimizer/powell fit and should be labeled as such, not reported as exact zeros.","section":"Table 2 and Table 4"},{"comment":"The sentence 'The direct model-free mean within-cell variance cross-check is computed and deposited with the bootstrap intervals in the v2 finalization' is ambiguous. If the value already exists in the repository, report it; if not, state clearly that it has not yet been computed.","section":"Section 5.2"},{"comment":"The two secondary corpora are described in the Data section but no decomposition is run on them. This is fine as a preview, but the section would be clearer if it explicitly said 'these corpora are not analyzed in this version' at the start of each subsection, rather than in a shared parenthetical.","section":"Section 3.2-3.3"},{"comment":"Reference [17] is missing author names. Some references [18]-[22] are to the author's own preprints and industry reports; please mark which items are peer-reviewed, since Section 2 currently mixes them with peer-reviewed literature without distinction.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is transparent and well-structured, and the structural argument is sound. My main concern is the gap between the provisional status of the estimates (explicitly conceded in Limitation 1 and Limitation 2) and the unqualified headline claims in the abstract and discussion. This is not an internal inconsistency, but it is a load-bearing gap: the practical conclusion depends on a Gaussian variance partition of a 91.9%-neutral outcome with no confidence intervals and a deferred model-free validation. I would encourage the editor to send this back for a revision that either supplies the missing bootstrap intervals and non-Gaussian sensitivity analysis, or visibly repositions the headline numbers as conditional on those pending checks."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a competent, honest application of standard generalizability theory to a real measurement problem, and the main practical conclusion — that repeats past five are the least efficient way to spend query budget — survives even if the outcome is misspecified. The paper deserves a serious referee, not a desk reject.\n\nWhat's actually new: a fully crossed four-facet decomposition (resampling, paraphrase, model, language) of response-level brand sentiment, with a decision-study allocation rule that says buy languages first, models second, paraphrases third, repeats last. The structural argument in Section 4.2 is the strongest part: the resampling term is divided by the full query count while the language term is divided by n_L only, so for any positive components the marginal value of a repeat decays faster. That does not depend on the fitted numbers, and the fitted numbers confirm it with a wide margin.\n\nThe paper also does some things well. The limitations section is unusually candid: it names the deferred cluster-bootstrap CIs, the missing model-free within-cell check, the low temperature, the grounding confound, and the zero-inflated outcome. The code is shipped, the REML fits are reproducible from the appendix, and the arithmetic in Tables 2–6 checks out. The self-citations are to prior work that is directly relevant, not padding.\n\nSoft spots, in proportion. The stress-test note has a real point: a Gaussian linear mixed model on an outcome where 91.9% of responses are exactly zero is fitting a mixture. The 26.5% language share and the 1.5% brand share could shift if the outcome were modeled as a two-part or ordinal process, because language may shift the probability of a neutral response rather than the sentiment of non-neutral ones. The paper acknowledges this as a \"first-order approximation,\" but the abstract states the split without that caveat, and the decision-study tables inherit the assumption. That said, this is not fatal to the main claim. The ranking reliability of a single answer is near zero under any reasonable model, and the repeats-last ordering is structural.\n\nOther concerns are minor or forward-looking. The CIs are deferred, so the point estimates have no interval guarantees. The response-level table is \"available on reasonable request\" rather than deposited, though the aggregate data and code are public. The model facet confounds model with retrieval mode. These are all named in the paper.\n\nWho this is for: anyone doing LLM brand tracking or reproducibility studies on LLM outputs. It is a useful empirical data point even now, and the methodological template is sound enough to build on.\n\nMy recommendation: send to peer review. The referee should push for the cluster-bootstrap intervals, the two-part or ordinal refit, and public release of the response-level table before full acceptance — but the core structural argument is in good shape and the paper is worth the referee time.","headline":"A transparent, well-scoped application of generalizability theory to LLM brand measurement, where the structural case against buying repeats is solid even though the headline variance split rests on a Gaussian fit to a 91.9%-neutral outcome.","tokens_in":14400,"tokens_out":1976,"would_cite":true,"duration_ms":23841,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Query language drives 26.5% of the variance in LLM brand answers, while brand identity drives only 1.5%.","keywords":["generalizability theory","variance components","intraclass correlation","large language models","measurement reproducibility","sampling design","brand measurement","query language"],"falsifier":"Recompute M1–M3 on the same 12,933 responses with a binary or ordinal recommendation outcome (brand named, or its rank) instead of sentiment polarity. If brand identity's ICC rises well above 0.0146 and language's share falls below 26.5%, the paper's central quantitative claim is an artifact of the near-degenerate sentiment outcome. Alternatively, fit a zero-inflated model; if the language variance component drops materially, the Gaussian assumption is the culprit.","tokens_in":13447,"feed_emoji":"🌐","tokens_out":3895,"duration_ms":35689,"temperature":0.7,"pith_summary":"The paper tries to establish where the run-to-run movement in LLM brand answers comes from, separating four sources: resampling the same prompt, paraphrasing the prompt, switching models, and asking in a different language. On its corpus, query language is by far the largest systematic source (26.5% of single-response variance) while brand identity is only 1.5% (ICC 0.0146), so one answer from one prompt in one language carries almost no information about where a brand ranks. The paper then uses the variance components to derive a spending rule: for a fixed query budget, adding languages and models reduces ranking error far more than adding more repeats, and repeats beyond five are the least efficient use of a query. If true, the common practice of averaging five repetitions of the same prompt is the wrong place to spend measurement effort.","feed_headline":"Language drives 26.5% of LLM brand-answer variance; brand just 1.5%","feed_subtitle":"A single AI answer can barely rank brands; spreading queries across languages and models beats repeating prompts.","key_machinery":"The crossed random-effects variance-components model (generalizability theory) with brand as object of measurement and language, model, and prompt as crossed random facets, plus a cell-level random intercept on the replicated subset so the residual becomes pure within-prompt resampling. The decision-study formula divides each variance component by the number of levels sampled for the facets it involves; because resampling is divided by the full query count while language is divided only by the language count, repeats have the fastest-decaying marginal value. The fitted components turn that structure into an explicit allocation rule.","core_discovery":"On a fully crossed corpus of 12,933 responses about 20 brands in 8 languages from 3 LLMs, the paper fits a crossed random-effects model partitioning the variance of a single brand-sentiment response into brand, language, model, prompt, interactions, and residual. The central result: query language accounts for 26.5% of the variance of one response against 1.5% for brand identity, and after isolating pure resampling on the replicated subset, within-prompt resampling is 34.8% while the brand-in-context interaction is 29.6%. Object-by-facet, only brand-by-language is sizable (8.6%); brand-by-model and brand-by-prompt are zero. Feeding these components into the decision-study equation shows bran","pith_inferences":["The near-degenerate outcome (91.9% neutral) likely suppresses brand signal; a recommendation-indicator refit may raise ICCs but probably preserves the facet ordering, since the ordering is structural.","The 26.5% language share is measured on sentiment polarity; on a binary 'is the brand named' outcome the language share could shrink or grow, and that is a direct testable extension from the same stored responses.","A zero-inflated or ordinal model could change the variance split; if the split holds under that model, the conclusion is much stronger.","The allocation rule has an immediate practical corollary for anyone tracking AI visibility: spend the next block of queries on a new language, not a sixth repeat."],"forward_implications":["If true, any brand ranking built from a single LLM answer (or a few repeats in one language) is essentially noise; brand identity explains only 1.5% of single-response variance.","Reliable brand measurement requires cross-language and cross-model breadth: 8 languages, 3 models, 1 paraphrase, 10 repeats reach Eρ² ≈ 0.13, whereas 20 repeats in one language/prompt reach only 0.02.","The field-standard 'repeat the prompt five times and average' convention is the least efficient allocation of a query budget; paraphrase breadth (15 paraphrases, 1 repeat) beats repeat depth (5 paraphrases, 5 repeats) at lower cost.","The reliability ceiling is low (~0.36) for this sentiment outcome; a less degenerate outcome such as a recommendation indicator is the natural next test.","The structural ordering—repeats saturate first—holds for any positive variance components, so the conclusion generalizes beyond the fitted numbers."],"fun_headline_variants":["Language drives 26.5% of LLM brand noise; brand just 1.5%","Repeating prompts won't fix LLM brand scores; add languages","LLM brand answers: language variance 26.5%, brand 1.5%","To rank brands with LLMs, vary languages not repeats","AI brand scores: language noise 26.5%, brand signal 1.5%"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole variance partition assumes the sentiment scores, 91.9% of which are exactly zero, can be treated as Gaussian continuous measurements; if a zero-inflated or ordinal model gives a different decomposition, the 26.5%/1.5%/34.8% splits and the allocation rule could change.","fun_headline_variants_meta":{"raw":{"variants":["Language drives 26.5% of LLM brand noise; brand just 1.5%","Repeating prompts won't fix LLM brand scores; add languages","LLM brand answers: language variance 26.5%, brand 1.5%","To rank brands with LLMs, vary languages not repeats","AI brand scores: language noise 26.5%, brand signal 1.5%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000286,"raw_usage":{"total_tokens":1621,"prompt_tokens":946,"completion_tokens":675,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":690,"completion_tokens_details":{"reasoning_tokens":569}},"tokens_in":690,"tokens_out":675,"duration_ms":6736,"temperature":1.0,"reasoning_tokens":569,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T05:34:56.187807+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute M1–M3 on the same 12,933 responses with a binary or ordinal recommendation outcome (brand named, or its rank) instead of sentiment polarity. If brand identity's ICC rises well above 0.0146 and language's share falls below 26.5%, the paper's central quantitative claim is an artifact of the near-degenerate sentiment outcome. Alternatively, fit a zero-inflated model; if the language variance component drops materially, the Gaussian assumption is the culprit.","supporting_citations":[],"review_version":1}