REVIEW 1 major objections 5 minor 27 references
The paper claims that forced-binary next-token probabilities from an instruction-tuned LLM are near-deterministic on 17 of 18 tested item pairs, so the resulting question-order (QQ) verdicts cannot identify a response mechanism.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 18:40 UTC pith:M3WBF3C2
load-bearing objection The theory is real and the pilot is honest, but the general health-check recommendation outruns the evidence from one model. the 1 major comments →
Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Central claim: the QQ equality—that the probability of disagreeing answers is order-invariant—holds for a broad family of classical and projective mechanisms, so a QQ-satisfied verdict cannot by itself identify a response mechanism. The audit criterion |qQQ| ≤ OSS (OSS the marginal order sensitivity) separates order effects from certified contextuality. Empirically, a pilot audit of an instruction-tuned model passed all health gates yet produced near-deterministic distributions on 17/18 direct-framing and 7/8 persona-framing items, label switching changed four verdicts, and no item was certified contextually. Hence observed QQ outcomes do not identify a mechanism.
What carries the argument
qQQ is the statistic: disagreement probability in order AB minus that in order BA; qQQ = 0 is the QQ equality. It is paired with OSS, the total order sensitivity of the yes-marginals; |qQQ| ≤ OSS is the audit form of the rank-2 Contextuality-by-Default criterion, separating order sensitivity, QQ imbalance, and residual contextuality. The measurement engine is a multi-turn forced-branch protocol: each first answer is committed to the history, both branches are evaluated, and order-conditioned joints are rebuilt from branch-weighted next-token log-probabilities under counterbalanced A/B label mappings. Pre-specified health gates, worst-case envelopes for unassigned mass, and a saturation diagn
Load-bearing premise
The practical recommendation—that saturation screening should be a standard health check whenever next-token probabilities are treated as survey responses—assumes that one open-weight instruction-tuned model, one chat template, one forced-binary label scheme, and 18 English items are representative of instruction-tuned LLM measurement in general; the paper's own limitation section disclaims cross-model generalization.
What would settle it
Run the same frozen audit protocol on several instruction-tuned models from different training pipelines, plus their non-instruction-tuned base checkpoints, on the same 18 item pairs. If a second instruction-tuned model produces dispersed (non-saturated) binary-conditioned distributions on a majority of items while passing all pre-specified health gates, the paper's general lesson—that saturation screening must precede interpretation of next-token-as-response audits—loses force. If a base checkpoint also saturates, the paper's candidate explanation that instruction-following post-training driv
If this is right
- A QQ 'satisfied' verdict can no longer be read as evidence for a projective or quantum mechanism, because a purely classical symmetric repetition mechanism satisfies the equality exactly.
- A QQ 'violated' verdict is only diagnostic when it exceeds the OSS allowance: violations with |qQQ| ≤ OSS remain compatible with noncontextual direct influences and do not certify contextuality.
- Forced-choice audits that skip a saturation diagnostic can mistake deterministic flips and label-assignment effects for genuine response mechanisms; the pilot shows this can happen even when all health gates pass.
- Label counterbalancing is required rather than optional: in the pilot, swapping the A/B assignment changed the mapping-specific QQ verdict on four of thirteen items.
- Distribution-level LLM audits should treat pre-specified saturation screening and explicit tracking of unassigned probability mass as prerequisites before interpreting structural criteria like QQ or contextuality.
Where Pith is reading between the lines
- Beyond the paper: the saturation finding, if it replicates, would put pressure on a broad class of LLM-as-survey-respondent studies that read forced-choice next-token probabilities as response distributions; the paper itself explicitly limits its empirical claim to the tested configuration.
- Beyond the paper: a natural testable extension is to run the same audit on base (non-instruction-tuned) checkpoints of the same model family; if saturation disappears there, instruction-following post-training—not forced-binary interfaces in general—would be implicated as the cause.
- Beyond the paper: the theory implies that ensemble or persona-mixture audits can restore dispersion, but only if mixture weights are order-matched and component-level diagnostics are reported; otherwise opposite-sign QQ violations across components could cancel in the pooled verdict.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops the quantum question (QQ) equality into an audit framework for sequential binary judgments of LLMs. The theoretical part characterizes mechanism families that satisfy QQ robustly (marginal-independent kernels, symmetric repetition, mixture-repetition families, order-matched mixing), proves a classical repetition mechanism that satisfies QQ identically, and translates the rank-2 Contextuality-by-Default criterion into the audit inequality |q_QQ| ≤ OSS. The methodological part introduces a multi-turn forced-branch protocol with counterbalanced label mappings, exact worst-case envelopes for unassigned mass, three health gates, and a pre-specified saturation diagnostic. A first-signal pilot on Qwen3-4B-Instruct-2507 under direct and persona framings finds the binary-conditioned distributions saturated for 17/18 (Track M) and 7/8 (Track P) item pairs, with label assignment changing mapping-specific QQ verdicts for four items and no item certified residually contextual. The paper concludes that next-token probabilities should not be treated as survey-response distributions without a dispersion check and recommends saturation screening as a standard health check.
Significance. If the theoretical results hold, they are a valuable contribution: they convert the QQ equality into a family of parameter-free mechanism characterizations, give an exact classical companion to the projective model, and provide a clean CbD translation that separates order sensitivity, QQ imbalance, and residual contextuality. The audit methodology is a genuine contribution — pre-specified gates, counterfactual branch commitment, worst-case mass envelopes, and semantic canonicalization are all carefully designed and transparently documented. The pilot is a useful cautionary case, backed by archived audit records, machine-validated algebra, and explicit limitation statements. The main weakness is external validity: the headline recommendation to make saturation screening a standard health check is inferred from a single model, one template, and a small item set, while the paper's own Section 7 disclaims cross-model generalization.
major comments (1)
- [§8 (Abstract, §5, §7)] The recommendation to treat saturation screening as 'a standard health check whenever next-token probabilities are treated as survey-response distributions' is an external-validity inference resting on a single configuration: Qwen3-4B-Instruct-2507, bf16, non-thinking, one chat template, one single-token A/B or Yes/No scheme, 18 English items (Track M) and 8 items (Track P). Section 7 explicitly disclaims cross-model generalization, and Section 6 lists model- and format-specific candidate explanations (RLHF-driven diversity reduction, the one-token 'Do not explain' instruction, item–prompt interactions). The empirical finding is therefore a case report, not evidence for a universal norm. Please either (a) scope the recommendation to the tested configuration and present it as a cautionary case study, or (b) add a small cross-model/cross-format validation (e.g., 2–3 models varying family o
minor comments (5)
- [§5 (Gate G3)] The claim 'all pre-specified health gates passed' should be qualified. The original G3 spot checks were run at batch size 64, and a later audit found batch-size-dependent generation distortion; the batch-one re-execution is a post hoc robustness check, not the originally pre-specified procedure. The recheck is appropriate and the verdicts are preserved, but the paper should explicitly state that the original G3 run was invalidated and that the pass is established by the retrospective re-execution.
- [Table 2] The table formatting contains column-wrap artifacts ('INDETERMINA TE', 'INDETERMINAT EY'). More substantively, s1-02 and s3-01 display Γ = 0.000 yet are classified INDETERMINATE; add a footnote explaining that the certified upper bound from the envelope intervals is positive at full precision even though displayed values round to zero.
- [Table 4 / §5 (finding 4)] The equal-weight pooling of map-1 and map-2 is justified by linearity and Proposition 6, but with label effects as large as those in Table 4 the pooled verdicts are mixture statistics rather than estimates of a single semantic response process. The paper does say this, but it would help to add one sentence in the Table 4 caption or the results text reminding readers that pooled values are not to be read as 'the' semantic effect.
- [Appendix A, Corollary 7 proof] The sign of OEA is opposite to E[A_AB]−E[A_BA] under the paper's convention (OEA = p_BA(Ay)−p_AB(Ay)). The absolute value hides this, but a reader may be momentarily confused. A one-line sign remark would make the translation more accessible.
- [Figure 2] The clipping of d=0 to 10^{-14} is mentioned in the body but should appear in the figure caption as well, since the caption is the natural place a reader will look when interpreting the left end of the CDF.
Circularity Check
No significant circularity: the theoretical derivations and empirical pilot are self-contained and reduce to no fitted inputs.
full rationale
The paper's derivation chain is self-contained. The theoretical results (Theorem 2, Theorem 3, Propositions 4-6, Corollary 7) are proved from explicit definitions in Appendix A, with Theorem 2 and Corollary 7 relying on external, independently published results (Wang-Busemeyer; Kujala-Dzhafarov) that are also numerically revalidated in the paper's own validation suite. The audit estimand P(y|yorn), the QQ statistic, OSS, and the saturation rule are definitions and pre-specified decision procedures, not parameters fitted to the pilot's verdicts. The pilot's saturation finding (17/18 and 7/8 items near-deterministic) is an empirical measurement made under pre-specified thresholds; it is not entailed by the saturation definition, since the definition alone does not predict that LLM positions will exceed 0.99. The label-counterbalancing effects are measured per-mapping quantities, and the pooled estimates are computed by the pre-specified equal-weight estimand rather than chosen to produce a desired QQ verdict. The paper explicitly disclaims cross-model generality in Section 7, so the concern that the 'standard health check' recommendation is over-broad is a limitation about external validity, not a circularity in the argument. No step in the paper reduces by construction to its own inputs, and no load-bearing appeal is made to the author's own prior work.
Axiom & Free-Parameter Ledger
free parameters (3)
- QQ practical-equivalence band εQQ =
0.02
- Saturation threshold on max(p, 1−p) =
0.99
- Gate G3 spot-check budget =
n=200, α=0.01 Bonferroni per item-order
axioms (7)
- domain assumption The Wang–Busemeyer projective sequential measurement model is the reference mechanism and predicts q_QQ = 0 for all states and binary projectors (Theorem 2).
- standard math The Kujala–Dzhafarov noncontextuality criterion for rank-2 cyclic systems with binary variables is correct and applicable to the two-order system.
- domain assumption Next-token probabilities of an autoregressive LLM at a forced answer position, renormalized to the two label tokens, can represent a survey-response distribution.
- domain assumption The sequential joint factorizes as pAB(α,β) = pA(α)KAB(β|α) under committed forced branching.
- domain assumption Deterministic logprob readout matches temperature-1 sampling from the same distribution.
- domain assumption Qwen3-4B-Instruct-2507 in bf16, non-thinking mode is informative about instruction-tuned forced-choice behavior in general.
- standard math Clopper–Pearson intervals give exact binomial confidence bounds (Bonferroni-corrected).
invented entities (3)
-
OSS (order-sensitivity score) = |OEA| + |OEB|
independent evidence
-
Γ = |q_QQ| − OSS (residual-contextuality certificate)
independent evidence
-
Binary-conditioned audit estimand P(y|yorn)
independent evidence
read the original abstract
Question-order effects in human survey data have been reported to approximately satisfy the QQ (quantum question) equality, a parameter-free prediction of the standard projective quantum question-order model. We develop this equality into an audit framework for sequential binary judgments of autoregressive large language models (LLMs). Theoretically, we characterize mechanism families that satisfy QQ robustly, show that classical repetition can reproduce the equality exactly, and combine QQ with the rank-2 Contextuality-by-Default criterion through $|q_{QQ}| \le \mathrm{OSS}$. This separates order sensitivity, QQ imbalance, and residual contextuality rather than treating them as interchangeable signatures. Methodologically, we introduce a committed multi-turn forced-branch protocol that reconstructs order-conditioned joint distributions from next-token log-probabilities under counterbalanced label mappings and pre-specified health gates. A first-signal pilot on an open-weight instruction-tuned model reveals the central measurement problem. Although all pre-specified health gates passed, the binary-conditioned distributions were near-deterministic for 17 of 18 item pairs under the direct-evaluation framing and 7 of 8 under the persona framing. Label assignment materially changed several mapping-specific QQ verdicts, and no item was certified as residually contextual. Thus, under the tested conditions, the observed QQ outcomes did not uniquely identify a response mechanism in the presence of a saturated and label-sensitive measurement interface. The main implication is methodological: next-token probabilities should not be interpreted as survey-response distributions without first establishing adequate dispersion. We therefore argue that saturation screening and label counterbalancing should precede structural interpretation in distribution-level audits of LLM judgments.
Figures
Reference graph
Works this paper leans on
-
[1]
David W. Moore. Measuring New Types of Question-Order Effects: Additive and Subtractive. Public Opinion Quarterly, 66(1):80–91, 2002. doi: 10.1086/338631
-
[2]
Zheng Wang and Jerome R. Busemeyer. A Quantum Question Order Model Supported by Empirical Tests of an A Priori and Precise Prediction.Topics in Cognitive Science, 5(4): 689–710, 2013. doi: 10.1111/tops.12040. URL https://onlinelibrary.wiley.com/doi/abs/ 10.1111/tops.12040
-
[3]
Zheng Wang, Tyler Solloway, Richard M. Shiffrin, and Jerome R. Busemeyer. Context Effects Produced by Question Orders Reveal Quantum Nature of Human Judgments.Proceedings of the National Academy of Sciences, 111(26):9431–9436, 2014. doi: 10.1073/pnas.1407756111
-
[4]
Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkar, and Graham Neubig. Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey Design.Transactions of the Association for Computational Linguistics, 12:1011–1026, 2024. doi: 10.1162/tacl a 00685. URLhttps://aclanthology.org/2024.tacl-1.56/
doi:10.1162/tacl 2024
-
[5]
Prompt Perturbations Reveal Human- Like Biases in Large Language Model Survey Responses
Jens Rupprecht, Georg Ahnert, and Markus Strohmaier. Prompt Perturbations Reveal Human- Like Biases in Large Language Model Survey Responses. InProceedings of the Seventh Workshop on Natural Language Processing and Computational Social Science (NLP+CSS), pages 1–21, San Diego, CA, 2026. Association for Computational Linguistics
2026
-
[6]
CAPE: Context-Aware Personality Evaluation Framework for Large Language Models
Jivnesh Sandhan, Fei Cheng, Tushar Sandhan, and Yugo Murawaki. CAPE: Context-Aware Personality Evaluation Framework for Large Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 10648–10662. Association for Computational Linguistics, 2025
2025
-
[7]
World Scientific, Singapore, 2023
Aleksandr Lebedev and Andrei Khrennikov.Quantum-like Modeling of the Order Effect in Decision Making: POVM Viewpoint on the Wang–Busemeyer QQ-equality, pages 123–128. World Scientific, Singapore, 2023. doi: 10.1142/9789811275999 0010. 21
-
[8]
Clopper and Egon S
Charles J. Clopper and Egon S. Pearson. The Use of Confidence or Fiducial Limits Illustrated in the Case of the Binomial.Biometrika, 26(4):404–413, 1934
1934
-
[9]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 Technical Report. arXiv preprint arXiv:2505.09388, 2025
Pith/arXiv arXiv 2025
-
[10]
Masanao Ozawa and Andrei Khrennikov. Modeling Combination of Question Order Ef- fect, Response Replicability Effect, and QQ-equality with Quantum Instruments.Jour- nal of Mathematical Psychology, 100:102491, 2021. ISSN 0022-2496. doi: https://doi.org/ 10.1016/j.jmp.2020.102491. URL https://www.sciencedirect.com/science/article/pii/ S0022249620301152
arXiv 2021
-
[11]
Dzhafarov, Janne V
Ehtibar N. Dzhafarov, Janne V. Kujala, and Victor H. Cervantes. Contextuality-by-Default: A Brief Overview of Ideas, Concepts, and Terminology. In Harald Atmanspacher, Thomas Filk, and Emmanuel Pothos, editors,Quantum Interaction, pages 12–23, Cham, 2016. Springer International Publishing
2016
-
[12]
Kujala and Ehtibar N
Janne V. Kujala and Ehtibar N. Dzhafarov. Proof of a Conjecture on Contextuality in Cyclic Systems with Binary Variables.Foundations of Physics, 46(3):282–299, Mar 2016. ISSN 1572-
2016
-
[13]
Failure of Contextual Invariance in Large Language Models, 2026
Sagar Kumar, Ariel Flint, Luca Maria Aiello, and Andrea Baronchelli. Failure of Contextual Invariance in Large Language Models, 2026. URLhttps://arxiv.org/abs/2603.23485
Pith/arXiv arXiv 2026
-
[14]
Thanh Luong Tuan. Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study, 2026. URLhttps://arxiv.org/abs/2607.11948
Pith/arXiv arXiv 2026
-
[15]
Pegah Imannezhad, Emmanuel M. Pothos, and Andy J. Wills. Divergent Patterns of Proba- bilistic Reasoning in Humans and GPT-5.Frontiers in Psychology, 17:1782184, 2026. ISSN 1664-1078. doi: 10.3389/fpsyg.2026.1782184. URL https://www.frontiersin.org/journals/ psychology/articles/10.3389/fpsyg.2026.1782184
arXiv 2026
-
[16]
Diederik Aerts, Jonito Aerts Argu¨ elles, Lester Beltran, Suzette Geriente, Roberto Leporini, Massimiliano Sassoli de Bianchi, and Sandro Sozzo. Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Human and Artificial Cognition.Entropy, 28(6):622, June 2026. ISSN 1099-4300. doi: 10.3390/e28060622. URL http://dx.doi.org/1...
-
[17]
Kin Ian Lo, Mehrnoosh Sadrzadeh, and Shane Mansfield. Quantum-like Contextuality in Large Language Models.Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences, 481(2319):20240399, 08 2025. ISSN 1364-5021. doi: 10.1098/rspa.2024.0399. URL https://doi.org/10.1098/rspa.2024.0399
arXiv 2025
-
[18]
Large Language Models Are Not Robust Multiple Choice Selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. Large Language Models Are Not Robust Multiple Choice Selectors. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=shr9PXz7T0
2024
-
[19]
Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions
Pouya Pezeshkpour and Estevam Hruschka. Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions. InFindings of the Association for Computational Linguistics: NAACL 2024, pages 2006–2017, 2024. 22
2024
-
[20]
My Answer is C
Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul R¨ ottger, Frauke Kreuter, Dirk Hovy, and Barbara Plank. “My Answer is C”: First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models. InFindings of the Association for Com- putational Linguistics: ACL 2024, pages 7407–7416. Association for Computational Linguistics, 2024
2024
-
[21]
doi: https://doi.org/10.1002/andp.19504430510
Gerhart L¨ uders.¨Uber die Zustands¨ anderung durch den Meßprozess.Annalen der Physik, 443 (5-8):322–328, 1950. doi: https://doi.org/10.1002/andp.19504430510
-
[22]
Understanding the Effects of RLHF on LLM Generalisation and Diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. Understanding the Effects of RLHF on LLM Generalisation and Diversity. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=PXD3FAVHJT
2024
-
[23]
Does Writing with Language Models Reduce Content Diversity? InThe Twelfth International Conference on Learning Representations, 2024
Vishakh Padmakumar and He He. Does Writing with Language Models Reduce Content Diversity? InThe Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Feiz5HtCD0
2024
-
[24]
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empiri...
2023
-
[25]
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec,...
Pith/arXiv arXiv 2022
-
[2023]
doi: 10.18653/v1/2023.emnlp-main.330
Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.330. URL https://aclanthology.org/2023.emnlp-main.330/
-
[9516]
URL https://doi.org/10.1007/s10701-015-9964-8
doi: 10.1007/s10701-015-9964-8. URL https://doi.org/10.1007/s10701-015-9964-8
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.