{"id":"365f7941-6488-4606-9d2b-505274a3fe4e","arxiv_id":"2608.10503","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A fully crossed factorial experiment on exact token probability distributions, combined with a distributional ANOVA, measures LLM ethnocentrism without text-sampling noise.","lead":"This paper introduces a way to measure an AI model's attitudes and biases by reading its exact next-word probabilities on a survey, instead of sampling many text responses. The approach aims to make AI behavior tests more reliable and to isolate whether a bias is a general trait, an effect of the prompt context, or an interaction between the two.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Effect distributions' SD, SNR, and dPD depend on the arbitrary comonotone coupling; only the expectations are coupling-invariant, so the reported 'exact' directional strengths are not model properties.","rationale":"The reader's weakest-assumption diagnosis is correct and matches my own read: the only quantities with a coupling-independent meaning are the expectations that Theorem 3.1 governs. The paper's headline framing, however, extends 'exact' to the full effect distributions, and the case-study conclusions lean on dPD and SNR values that vary with the coupling. This does not invalidate the core distributional ANOVA idea or the theorem, but it does require a sensitivity analysis and softened language in the reported directional claims. I did not find a more load-bearing internal inconsistency: the expectation-level proofs appear sound, the convolution and paired-comonotone computations are deterministic, and the case study is reproducible in principle from the provided PMFs. The Consensus equation in the main text may also contain a dimensionally inconsistent squared-norm typo, but that is a separate, readily correctable presentation error and not the central concern. Since the reader already concluded CONDITIONAL and my concern reinforces that conclusion rather than moving it, I recommend UNCHANGED.","tokens_in":34100,"tokens_out":11587,"duration_ms":110327,"concrete_test":"Recompute all rows of Table 6 (and Table 9 Panel B) with the same marginal PMFs but replace the shared quantile U* in Def. A.10 with independent uniforms per term (the product coupling), keeping every marginal fixed; then recompute dPD and SNR via Eq. (82)-(85). Also compute the Frechet bounds for P(E>0) over all couplings for each reported interaction. If the comonotone dPD values (e.g., Gemma3 x USA 0.63, Llama3.3 x USA 0.58) move by more than roundoff or cross the 0.5 chance boundary, the directional claims are coupling artifacts and must be re-reported as conditional on the chosen coupling.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central methodological guarantee (Theorem 3.1; App. A.4.5, A.4.6) is that the expectations of the effect distributions recover the unique Hoeffding/ANOVA decomposition. This is correct: Def. A.8/A.10 couple random variables with the same marginals, so E[EU] = µ_U(λ_U) regardless of the coupling. However, the paper then reports SD, SNR, and dPD as exact properties of the model's country-of-origin bias. These dispersion and directionality summaries are not coupling-invariant. The comonotone coupling used in Eq. (42) and Eq. (69) is one admissible pairing among infinitely many with identical marginals; under an independent or countermonotone coupling, P(E>0), Var(E), and therefore dPD and SNR change, while the marginals, and Theorem 3.1, remain unchanged. For example, the US ingroup dPD values 0.63 and 0.58 are presented as evidence of moderate directional favoritism, but they are artifacts of the chosen maximal-dependence pairing, not intrinsic features of the model's predictive distributions. The paper's own framing in App. A.4.5 ('minimizes variance') indicates a variance-reduction heuristic rather than an empirically grounded canonical pairing. The Limitations section does not flag this dependency. A coupling sensitivity analysis is therefore required before these directional strengths can legitimately be called exact model properties.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an exact-PMF framework for measuring LLM attitudes and biases, replacing Monte Carlo text sampling with direct token-level probability mass functions. It introduces a fully crossed factorial design, a multivariate ordinal Consensus metric, and a distributional ANOVA/Hoeffding decomposition that isolates baseline, main-effect, and interaction-effect distributions. A case study on the CETSCALE across five LLMs claims to expose country-of-origin interaction effects that aggregate benchmarks obscure, and an analysis of sampling cost shows that standard finite-sample estimators can flip the sign of small effects.","tokens_in":34394,"tokens_out":9011,"duration_ms":86020,"significance":"The core theoretical contribution—that the expectations of the construction's effect distributions recover the unique Hoeffding/ANOVA decomposition—is proved carefully in Appendix A.4 and appears correct. If the framework holds, it would give the NLP community a principled way to attribute behavioral differences to main effects versus interactions without sampling noise, using exact convolutions and closed-form contrast distributions. The detailed appendices, transparent case study, and explicit falsifiable predictions are strengths. However, two issues currently limit the paper: the published Consensus formula is dimensionally inconsistent, and the reported dispersion-based summaries (SD, SNR, dPD) depend on an arbitrary comonotone coupling rather than being intrinsic model properties. Both are fixable, so the result is not fundamentally unsound, but the exactness claims need to be qualified.","major_comments":[{"comment":"The multivariate Consensus as written divides a squared Euclidean distance by a linear distance: with dmax defined as a maximum distance on the scale, the argument 1 − ||y−µ||²/dmax can become negative, making the logarithm undefined. For K=17 and a 7-point scale, such negative arguments arise already for moderate deviations from the centroid, so the published formula cannot be what produced the values in Table 5. The normalization should presumably be by dmax², or the numerator should use the unsquared distance. Please correct the definition, restate the text, and recompute or confirm the affected Consensus values.","section":"3.3, Eq. (1); App. A.3, Eq. (12)"},{"comment":"Theorem 3.1 guarantees coupling-invariance only for expectations. The SD, SNR, and dPD values reported in Tables 2, 6, and 9 are computed under the comonotone (maximal-dependence) coupling, and any other coupling with the same marginals—independent, countermonotone, or otherwise—preserves the expectation-level theorem but changes these dispersion and directionality summaries. The dPD values of 0.63 and 0.58 cited as evidence of moderate US ingroup favoritism are therefore not intrinsic properties of the models' predictive distributions. The paper should report a coupling-sensitivity analysis or explicitly qualify every dispersion-based summary as conditional on the comonotone pairing.","section":"4.2, Table 6, Table 9; App. A.4.5–A.4.6, Def. A.8/A.10, Eqs. (42), (69)"},{"comment":"The comparison between the aggregate Target-Country main effect and the Model×Target interaction is presented as an empirical demonstration that aggregate benchmarks are 'directionally incorrect.' Because main effects and interactions in a fully crossed ANOVA decomposition are orthogonal by construction, the sign reversal between Panel A and Panel B is a mathematical necessity, not a data-dependent discovery. The text should state this explicitly; as written, it overstates the empirical content of the comparison.","section":"4.2, Table 9"}],"minor_comments":[{"comment":"The statement that |Vval| = |Y| contradicts the preceding description and Section 3.2, where multiple token surface forms (e.g., \" 7\" and \"7\") map to the same ordinal value. The mapping φ is surjective but not injective when tokenizer variants exist; please remove the cardinality equality.","section":"App. A.1 vs. Sec. 3.2"},{"comment":"The phrase \"As demonstrated in 3\" should read \"As demonstrated in Section 3.\"","section":"Sec. 2"},{"comment":"Several entropy values exceed log2(7) ≈ 2.807 (e.g., 7.315, 35.519), so the entropy is evidently summed over the 17 items. The table caption and the surrounding text should state this explicitly, because the Consensus values are not summed and the two metrics are otherwise not comparable.","section":"Table 5"},{"comment":"The main text says Theorem 3.1 is proved in App. A.9, but App. A.9 concerns sampling estimators; the actual proof appears in App. A.4.4–A.4.6. Please correct the cross-reference.","section":"Sec. 3.4 / App. A.4"},{"comment":"The display for dmax is typeset ambiguously (the exponent on (y_max−y_min) is unclear), and the phrase \"maximum diagonal distance on the Likert scale\" is imprecise. Please write the formula explicitly and clarify that dmax is a distance in the K-dimensional response space.","section":"Eq. (12)"},{"comment":"The text compares SNR to Cohen's d but uses different definitions; please clarify that SNR is not Cohen's d and that the heuristic thresholds are only descriptive references, to avoid potential misinterpretation.","section":"App. A.4.11"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically ambitious and the expectation-level theorem is sound, but the two load-bearing issues—the Consensus formula's dimensional error and the coupling-dependence of the reported dispersion summaries—prevent acceptance in the current form. The paper would also benefit from a statement about the availability of code and exact model outputs, since the empirical claims are not reproducible from the text alone. The citation list contains a large share of self-citations and several forthcoming or 2026-dated references; this is not disqualifying, but the editor may wish to verify their availability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on 2608.10503. The distributional ANOVA core is genuinely interesting and mostly sound, but the paper overstates the exactness of its dispersion summaries, and the multivariate Consensus equation is dimensionally wrong as printed.\n\nWhat's new: the combination of fully crossed factorial prompts, exact next-token PMF extraction, and a distributional Hoeffding decomposition where each main and interaction effect is a probability distribution instead of a point estimate. Theorem 3.1, which says the expectations of those effect distributions recover the classical unique decomposition, is proven correctly in the appendix: I checked the Möbius-inversion argument and it's solid. The multivariate extension of Tastle-Wierman consensus is a reasonable idea, and the CETSCALE case study across five LLMs is a good proof of concept. The prompt-framing extension is a nice touch: it shows the framework can formally quantify wording sensitivity instead of treating it as a nuisance.\n\nThe soft spots are real but fixable. First, Eq. (1) divides a squared distance by a distance; for K=17 items, the log argument can go negative, so the formula as written is not a valid probability-weighted log. They presumably meant dmax² in the denominator, but as printed it's wrong. Second—and this is the more substantive issue—the reported SD, SNR, and dPD of the effect distributions depend on the comonotone coupling used to define paired contrasts. The expectations are coupling-invariant (that's Theorem 3.1), but the dispersion and direction metrics are not. The paper presents dPD=0.63 for the US ingroup interaction as an exact model property, yet under a different coupling it would change. The appendix even says the comonotone choice 'minimizes variance'—a sensible heuristic, but not a canonical property of the model. There's no sensitivity analysis and no caveat in the Limitations. That makes some of the specific directional findings overclaimed, even though the mean-shift results stand.\n\nThe math and data are otherwise in good shape: the appendices are careful, the PMF computations are deterministic, and the citation pattern looks normal—the self-citations are to relevant prior work. The paper is for researchers doing LLM bias audits or psychometric-style evaluations who want more causal structure than aggregate benchmarks provide. It deserves a serious referee: send it to review, but require (1) fixing Eq. (1), (2) a coupling sensitivity analysis, or at minimum an explicit statement that dPD/SNR are properties of the chosen paired contrast, and (3) ideally a code/data release. With those changes, this could be a solid methods contribution.","headline":"Worth a serious referee: the distributional ANOVA core is sound, but the multivariate Consensus formula is dimensionally wrong as written and the reported dispersion/direction 'exactness' depends on an arbitrary unacknowledged coupling choice.","tokens_in":34875,"tokens_out":3909,"would_cite":false,"duration_ms":35054,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Exact token-level probability distributions, run through a fully crossed factorial ANOVA, isolate causal LLM biases that aggregate benchmarks miss.","keywords":["LLM behavioral evaluation","token-level probability mass functions","fully crossed factorial design","distributional ANOVA","Hoeffding decomposition","multivariate consensus","country-of-origin bias","exact inference"],"falsifier":"Recompute the interaction summaries in Table 6 with independent or countermonotone pairing instead of the comonotone coupling; if the dPD and SNR of the US-model ingroup cells change materially, the reported directionality is a coupling artifact rather than a property of the model.","tokens_in":33925,"feed_emoji":"📊","tokens_out":8856,"duration_ms":80001,"temperature":0.7,"pith_summary":"This paper sets out to make LLM attitude and bias measurement exact rather than sampled. It replaces unstructured prompting with fully crossed factorial experiments and replaces Monte Carlo text generation with the model's own next-token probability mass functions, then processes those PMFs analytically. The central claim is Theorem 3.1: in any fully crossed design, the expectations of the isolated main-effect and interaction-effect distributions recover the classical unique Hoeffding/ANOVA decomposition, so a construct like country-of-origin bias can be measured as an interaction effect with baseline and main effects stripped away. If the claim holds, behavioral evaluation of LLMs becomes deterministic, free of sampling noise, and able to attribute a bias to baseline traits, contexts, or their interaction instead of to correlated prompt content. The paper demonstrates the pipeline on a five-model ethnocentrism case study in which aggregate benchmarks are directionally wrong for specific models.","feed_headline":"Exact token probabilities expose LLM bias without sampling noise","feed_subtitle":"One-pass exact PMFs separate baseline, main, and interaction effects, exposing biases aggregate benchmarks miss.","key_machinery":"The load-bearing object is the distributional Hoeffding/ANOVA decomposition built from paired contrasts. Each effect distribution is formed by drawing from two marginal PMFs through a shared uniform quantile, the comonotone coupling, which preserves the marginals while minimizing the variance of the difference; this is what makes the contrast's expectation equal to the classical ANOVA effect while retaining the full PMF. Around that core, discrete convolution propagates the item-level PMFs into an exact composite-score distribution, and a multivariate generalization of the Consensus metric supplies the ordinal-aware certainty measure that Shannon entropy lacks.","core_discovery":"The central discovery is that a fully crossed factorial experiment over exact token-level PMFs yields a distributional ANOVA whose expectations reproduce the classical unique Hoeffding decomposition. The paper builds a grand-mixture baseline and marginal-slice distributions, then defines main-effect and interaction-effect distributions as paired contrasts under a comonotone coupling, so the expectation of each effect equals the corresponding fixed-effects ANOVA parameter. This is what makes country-of-origin bias a well-defined interaction term: the US-developed models show positive own-country interactions (+3.21 for Gemma and +2.59 for Llama), while the aggregate target-country main effect would have hidden the sign reversals visible in the interaction table. The intended reading is that the pipeline is exact at the distributional level for means, with all aleatoric uncertainty propagated from tokens to the composite score.","pith_inferences":["Exactness is proven for expectations; the spread and direction metrics inherit the chosen comonotone coupling, so an editor would want replications to report whether dPD and SNR are stable under independent or countermonotone pairing.","The grand-mixture baseline weights every experimental condition equally, so a deployment-realistic baseline would need prevalence-weighted conditions; the theorem still applies to those weights.","The same decomposition could separate a model that always disfavors a demographic (main effect) from one that disfavors it only in specific contexts (interaction), which is exactly the gender-bias question the introduction poses.","The framework's exactness is scoped to constrained single-token responses; extending it to chain-of-thought would require an integrated distribution over latent multi-token paths, not just a convolution of item-level PMFs."],"forward_implications":["Small effect parameters that flip sign 18% of the time at N=10 under text sampling are recovered exactly in one forward pass, so subtle interactions can be measured without Monte Carlo noise.","Aggregate target-country effects can be directionally wrong for individual models: the paper finds a negative aggregate France effect but a positive French interaction for Ministral.","Confounders such as prompt framing can be added as crossed factors, turning prompt sensitivity into an isolated main effect and model-by-framing interaction instead of an uncontrolled critique.","The same Theorem 3.1 guarantee applies to any fully crossed design, so any ordinal instrument can be decomposed into baseline, main, and interaction effect distributions with the same machinery.","Exact convolution of item PMFs propagates all aleatoric uncertainty from tokens to the final composite score, so downstream comparisons carry the model's full response distribution rather than point estimates."],"supporting_citations":[{"why":"Supplies the 17-item CETSCALE Likert instrument and the historical human sample means used to anchor the case study.","marker":"Shimp and Sharma (1987)"},{"why":"Defines the univariate Consensus measure that the paper generalizes to a multivariate, PMF-based ordinal consensus metric.","marker":"Tastle and Wierman (2007)"},{"why":"Provides the Möbius inversion on the Boolean lattice used to derive the unique Hoeffding/ANOVA components in Theorem 3.1 and its corollaries.","marker":"Rota (1964)"},{"why":"Supplies the optimal-transport/comonotone coupling used for paired contrasts that define main- and interaction-effect distributions.","marker":"Villani et al. (2008)"},{"why":"Precedent for treating Likert items as conditionally independent given the prompt and for characterizing sampling-based evaluation noise that the exact-PMF method removes.","marker":"Wadi and Fredette (2025)"},{"why":"Source of the probability-of-direction concept adapted into the discrete dPD metric for discrete effect distributions.","marker":"Makowski et al. (2019)"}],"fun_headline_variants":["Exact token PMFs isolate LLM bias, end sampling noise","Factorial exact PMFs reveal LLM bias benchmarks miss","Distributional ANOVA on tokens separates bias causes","Noise-free token analysis uncovers LLM interaction bias","Exact token probabilities: analytic LLM bias isolation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported spreads, signal-to-noise ratios, and directional probabilities depend on the paper's chosen pairing of the compared distributions, and if that pairing is not canonical, those strengths and directions can change even though the effect means remain fixed.","fun_headline_variants_meta":{"raw":{"variants":["Exact token PMFs isolate LLM bias, end sampling noise","Factorial exact PMFs reveal LLM bias benchmarks miss","Distributional ANOVA on tokens separates bias causes","Noise-free token analysis uncovers LLM interaction bias","Exact token probabilities: analytic LLM bias isolation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000362,"raw_usage":{"total_tokens":1934,"prompt_tokens":904,"completion_tokens":1030,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":951}},"tokens_in":520,"tokens_out":1030,"duration_ms":9976,"temperature":1.0,"reasoning_tokens":951,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:19:04.425588+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the interaction summaries in Table 6 with independent or countermonotone pairing instead of the comonotone coupling; if the dPD and SNR of the US-model ingroup cells change materially, the reported directionality is a coupling artifact rather than a property of the model.","supporting_citations":[],"review_version":1}