Pith. sign in

REVIEW 1 major objections 5 minor 27 references

The paper claims that forced-binary next-token probabilities from an instruction-tuned LLM are near-deterministic on 17 of 18 tested item pairs, so the resulting question-order (QQ) verdicts cannot identify a response mechanism.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 18:40 UTC pith:M3WBF3C2

load-bearing objection The theory is real and the pilot is honest, but the general health-check recommendation outruns the evidence from one model. the 1 major comments →

arxiv 2607.17219 v2 pith:M3WBF3C2 submitted 2026-07-19 cs.CL cs.AIquant-phstat.ME

Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

classification cs.CL cs.AIquant-phstat.ME
keywords question-order effectsQQ equalityquantum question modellarge language modelsnext-token probabilitiescontextuality-by-defaultsaturation diagnosticforced-choice response measurement
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Human survey answers to paired questions show an empirical regularity: the probability that the two answers disagree is nearly the same regardless of which question is asked first—the QQ (quantum question) equality. This paper turns that equality into an audit for sequence-sensitive large language models (LLMs), asking whether forced-binary LLM judgments obey the same constraint and what a verdict could mean. It proves that QQ satisfaction is compatible with several distinct mechanisms, including a purely classical repetition rule that reproduces the equality exactly, so a 'satisfied' result alone never identifies how the model responds. Empirically, running the audit's multi-turn protocol on one open-weight instruction-tuned model produced near-deterministic answer distributions on 17 of 18 direct-framing items and 7 of 8 persona-framing items, even though every pre-specified health gate passed and label assignment changed several verdicts. The paper's conclusion: next-token probabilities should not be treated as survey-response distributions without first checking dispersion, and saturation screening plus label counterbalancing should precede structural interpretation in LLM judgment audits.

Core claim

Central claim: the QQ equality—that the probability of disagreeing answers is order-invariant—holds for a broad family of classical and projective mechanisms, so a QQ-satisfied verdict cannot by itself identify a response mechanism. The audit criterion |qQQ| ≤ OSS (OSS the marginal order sensitivity) separates order effects from certified contextuality. Empirically, a pilot audit of an instruction-tuned model passed all health gates yet produced near-deterministic distributions on 17/18 direct-framing and 7/8 persona-framing items, label switching changed four verdicts, and no item was certified contextually. Hence observed QQ outcomes do not identify a mechanism.

What carries the argument

qQQ is the statistic: disagreement probability in order AB minus that in order BA; qQQ = 0 is the QQ equality. It is paired with OSS, the total order sensitivity of the yes-marginals; |qQQ| ≤ OSS is the audit form of the rank-2 Contextuality-by-Default criterion, separating order sensitivity, QQ imbalance, and residual contextuality. The measurement engine is a multi-turn forced-branch protocol: each first answer is committed to the history, both branches are evaluated, and order-conditioned joints are rebuilt from branch-weighted next-token log-probabilities under counterbalanced A/B label mappings. Pre-specified health gates, worst-case envelopes for unassigned mass, and a saturation diagn

Load-bearing premise

The practical recommendation—that saturation screening should be a standard health check whenever next-token probabilities are treated as survey responses—assumes that one open-weight instruction-tuned model, one chat template, one forced-binary label scheme, and 18 English items are representative of instruction-tuned LLM measurement in general; the paper's own limitation section disclaims cross-model generalization.

What would settle it

Run the same frozen audit protocol on several instruction-tuned models from different training pipelines, plus their non-instruction-tuned base checkpoints, on the same 18 item pairs. If a second instruction-tuned model produces dispersed (non-saturated) binary-conditioned distributions on a majority of items while passing all pre-specified health gates, the paper's general lesson—that saturation screening must precede interpretation of next-token-as-response audits—loses force. If a base checkpoint also saturates, the paper's candidate explanation that instruction-following post-training driv

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A QQ 'satisfied' verdict can no longer be read as evidence for a projective or quantum mechanism, because a purely classical symmetric repetition mechanism satisfies the equality exactly.
  • A QQ 'violated' verdict is only diagnostic when it exceeds the OSS allowance: violations with |qQQ| ≤ OSS remain compatible with noncontextual direct influences and do not certify contextuality.
  • Forced-choice audits that skip a saturation diagnostic can mistake deterministic flips and label-assignment effects for genuine response mechanisms; the pilot shows this can happen even when all health gates pass.
  • Label counterbalancing is required rather than optional: in the pilot, swapping the A/B assignment changed the mapping-specific QQ verdict on four of thirteen items.
  • Distribution-level LLM audits should treat pre-specified saturation screening and explicit tracking of unassigned probability mass as prerequisites before interpreting structural criteria like QQ or contextuality.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the saturation finding, if it replicates, would put pressure on a broad class of LLM-as-survey-respondent studies that read forced-choice next-token probabilities as response distributions; the paper itself explicitly limits its empirical claim to the tested configuration.
  • Beyond the paper: a natural testable extension is to run the same audit on base (non-instruction-tuned) checkpoints of the same model family; if saturation disappears there, instruction-following post-training—not forced-binary interfaces in general—would be implicated as the cause.
  • Beyond the paper: the theory implies that ensemble or persona-mixture audits can restore dispersion, but only if mixture weights are order-matched and component-level diagnostics are reported; otherwise opposite-sign QQ violations across components could cancel in the pooled verdict.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. The paper develops the quantum question (QQ) equality into an audit framework for sequential binary judgments of LLMs. The theoretical part characterizes mechanism families that satisfy QQ robustly (marginal-independent kernels, symmetric repetition, mixture-repetition families, order-matched mixing), proves a classical repetition mechanism that satisfies QQ identically, and translates the rank-2 Contextuality-by-Default criterion into the audit inequality |q_QQ| ≤ OSS. The methodological part introduces a multi-turn forced-branch protocol with counterbalanced label mappings, exact worst-case envelopes for unassigned mass, three health gates, and a pre-specified saturation diagnostic. A first-signal pilot on Qwen3-4B-Instruct-2507 under direct and persona framings finds the binary-conditioned distributions saturated for 17/18 (Track M) and 7/8 (Track P) item pairs, with label assignment changing mapping-specific QQ verdicts for four items and no item certified residually contextual. The paper concludes that next-token probabilities should not be treated as survey-response distributions without a dispersion check and recommends saturation screening as a standard health check.

Significance. If the theoretical results hold, they are a valuable contribution: they convert the QQ equality into a family of parameter-free mechanism characterizations, give an exact classical companion to the projective model, and provide a clean CbD translation that separates order sensitivity, QQ imbalance, and residual contextuality. The audit methodology is a genuine contribution — pre-specified gates, counterfactual branch commitment, worst-case mass envelopes, and semantic canonicalization are all carefully designed and transparently documented. The pilot is a useful cautionary case, backed by archived audit records, machine-validated algebra, and explicit limitation statements. The main weakness is external validity: the headline recommendation to make saturation screening a standard health check is inferred from a single model, one template, and a small item set, while the paper's own Section 7 disclaims cross-model generalization.

major comments (1)
  1. [§8 (Abstract, §5, §7)] The recommendation to treat saturation screening as 'a standard health check whenever next-token probabilities are treated as survey-response distributions' is an external-validity inference resting on a single configuration: Qwen3-4B-Instruct-2507, bf16, non-thinking, one chat template, one single-token A/B or Yes/No scheme, 18 English items (Track M) and 8 items (Track P). Section 7 explicitly disclaims cross-model generalization, and Section 6 lists model- and format-specific candidate explanations (RLHF-driven diversity reduction, the one-token 'Do not explain' instruction, item–prompt interactions). The empirical finding is therefore a case report, not evidence for a universal norm. Please either (a) scope the recommendation to the tested configuration and present it as a cautionary case study, or (b) add a small cross-model/cross-format validation (e.g., 2–3 models varying family o
minor comments (5)
  1. [§5 (Gate G3)] The claim 'all pre-specified health gates passed' should be qualified. The original G3 spot checks were run at batch size 64, and a later audit found batch-size-dependent generation distortion; the batch-one re-execution is a post hoc robustness check, not the originally pre-specified procedure. The recheck is appropriate and the verdicts are preserved, but the paper should explicitly state that the original G3 run was invalidated and that the pass is established by the retrospective re-execution.
  2. [Table 2] The table formatting contains column-wrap artifacts ('INDETERMINA TE', 'INDETERMINAT EY'). More substantively, s1-02 and s3-01 display Γ = 0.000 yet are classified INDETERMINATE; add a footnote explaining that the certified upper bound from the envelope intervals is positive at full precision even though displayed values round to zero.
  3. [Table 4 / §5 (finding 4)] The equal-weight pooling of map-1 and map-2 is justified by linearity and Proposition 6, but with label effects as large as those in Table 4 the pooled verdicts are mixture statistics rather than estimates of a single semantic response process. The paper does say this, but it would help to add one sentence in the Table 4 caption or the results text reminding readers that pooled values are not to be read as 'the' semantic effect.
  4. [Appendix A, Corollary 7 proof] The sign of OEA is opposite to E[A_AB]−E[A_BA] under the paper's convention (OEA = p_BA(Ay)−p_AB(Ay)). The absolute value hides this, but a reader may be momentarily confused. A one-line sign remark would make the translation more accessible.
  5. [Figure 2] The clipping of d=0 to 10^{-14} is mentioned in the body but should appear in the figure caption as well, since the caption is the natural place a reader will look when interpreting the left end of the CDF.

Circularity Check

0 steps flagged

No significant circularity: the theoretical derivations and empirical pilot are self-contained and reduce to no fitted inputs.

full rationale

The paper's derivation chain is self-contained. The theoretical results (Theorem 2, Theorem 3, Propositions 4-6, Corollary 7) are proved from explicit definitions in Appendix A, with Theorem 2 and Corollary 7 relying on external, independently published results (Wang-Busemeyer; Kujala-Dzhafarov) that are also numerically revalidated in the paper's own validation suite. The audit estimand P(y|yorn), the QQ statistic, OSS, and the saturation rule are definitions and pre-specified decision procedures, not parameters fitted to the pilot's verdicts. The pilot's saturation finding (17/18 and 7/8 items near-deterministic) is an empirical measurement made under pre-specified thresholds; it is not entailed by the saturation definition, since the definition alone does not predict that LLM positions will exceed 0.99. The label-counterbalancing effects are measured per-mapping quantities, and the pooled estimates are computed by the pre-specified equal-weight estimand rather than chosen to produce a desired QQ verdict. The paper explicitly disclaims cross-model generality in Section 7, so the concern that the 'standard health check' recommendation is over-broad is a limitation about external validity, not a circularity in the argument. No step in the paper reduces by construction to its own inputs, and no load-bearing appeal is made to the author's own prior work.

Axiom & Free-Parameter Ledger

3 free parameters · 7 axioms · 3 invented entities

The central theoretical results (Theorem 3, Propositions 4–6, Corollary 7) introduce no free parameters and are hand-checkable algebra consistent with the cited external benchmarks. The audit's free parameters are pre-specified thresholds (εQQ, saturation cutoff, G3 budget) of the legitimate kind; none is fitted to pilot outcomes, and the saturation finding is threshold-insensitive. Load-bearing domain assumptions are the estimand's validity (explicitly flagged and tested by the paper), the committed-branch factorization, and the single-model generality premise — the last being the paper's own stated limitation. The invented quantities are audit statistics with direct falsifiable handles, not unexplained entities.

free parameters (3)
  • QQ practical-equivalence band εQQ = 0.02
    Pre-specified boundary separating SATISFIED from VIOLATED; anchored to partial-revision simulations in Wang et al. (2014) SI (|q| < 0.015), not fitted to pilot data. Affects verdict labeling, not the saturation finding.
  • Saturation threshold on max(p, 1−p) = 0.99
    Pre-specified cutoff for the saturation diagnostic; the author reports insensitivity across 0.95/0.975/0.99 (18/18, 18/18, 17/18 saturated in Track M), so it does not drive the main conclusion.
  • Gate G3 spot-check budget = n=200, α=0.01 Bonferroni per item-order
    Pre-specified sampling-consistency parameters; procedural rather than fitted. The method (Clopper–Pearson) was selected after observing Wilson-interval undercoverage in calibration, which is disclosed.
axioms (7)
  • domain assumption The Wang–Busemeyer projective sequential measurement model is the reference mechanism and predicts q_QQ = 0 for all states and binary projectors (Theorem 2).
    The audit framework's benchmark, inherited from prior literature [2,3]; the paper reproduces it numerically but does not re-derive it.
  • standard math The Kujala–Dzhafarov noncontextuality criterion for rank-2 cyclic systems with binary variables is correct and applicable to the two-order system.
    Corollary 7 and all Γ-certification verdicts depend on this external theorem [12]; the Appendix proof is an algebraic translation, not a proof of the criterion itself.
  • domain assumption Next-token probabilities of an autoregressive LLM at a forced answer position, renormalized to the two label tokens, can represent a survey-response distribution.
    The binary-conditioned estimand of §4. The paper explicitly flags this as the measurement assumption under test; the pilot's saturation finding shows it fails in the tested regime.
  • domain assumption The sequential joint factorizes as pAB(α,β) = pA(α)KAB(β|α) under committed forced branching.
    Equation (5) in §4; committing the first-answer token is the operational analogue of the projective state update. If a committed token diverges from the model's actual sequential conditioning, reconstructed joints are invalid.
  • domain assumption Deterministic logprob readout matches temperature-1 sampling from the same distribution.
    Gate G3 validates this on two items (Track M) and one item (Track P) via Clopper–Pearson spot checks; a post hoc batch-size distortion was found and rechecked at batch size one.
  • domain assumption Qwen3-4B-Instruct-2507 in bf16, non-thinking mode is informative about instruction-tuned forced-choice behavior in general.
    Underpins the general saturation-screening recommendation; §7 explicitly limits claims to the tested configuration, leaving this assumption open.
  • standard math Clopper–Pearson intervals give exact binomial confidence bounds (Bonferroni-corrected).
    Used in Gate G3 spot checks; standard statistical result.
invented entities (3)
  • OSS (order-sensitivity score) = |OEA| + |OEB| independent evidence
    purpose: CbD inconsistent-connectedness allowance in audit coordinates; enters the certification inequality |q_QQ| ≤ OSS.
    A defined statistic directly computable from measured marginals; its interpretive role is anchored externally by the Kujala–Dzhafarov criterion.
  • Γ = |q_QQ| − OSS (residual-contextuality certificate) independent evidence
    purpose: Three-way certification verdict (certified noncontextual / indeterminate / certified residual contextuality) via sound-but-conservative envelope bounds Γlower ≤ Γ ≤ Γupper.
    Computable from data; the sign of Γ is a falsifiable handle, and the paper reports the indeterminate cases honestly.
  • Binary-conditioned audit estimand P(y|yorn) independent evidence
    purpose: Survey-response proxy formed by renormalizing two next-token log-probabilities.
    Checked against trajectory-level sampling by Gate G3; the pilot's central finding is that this proxy can be degenerate while passing the checks.

pith-pipeline@v1.3.0-alltime-deepseek · 23441 in / 21510 out tokens · 191561 ms · 2026-08-01T18:40:20.962262+00:00 · methodology

0 comments
read the original abstract

Question-order effects in human survey data have been reported to approximately satisfy the QQ (quantum question) equality, a parameter-free prediction of the standard projective quantum question-order model. We develop this equality into an audit framework for sequential binary judgments of autoregressive large language models (LLMs). Theoretically, we characterize mechanism families that satisfy QQ robustly, show that classical repetition can reproduce the equality exactly, and combine QQ with the rank-2 Contextuality-by-Default criterion through $|q_{QQ}| \le \mathrm{OSS}$. This separates order sensitivity, QQ imbalance, and residual contextuality rather than treating them as interchangeable signatures. Methodologically, we introduce a committed multi-turn forced-branch protocol that reconstructs order-conditioned joint distributions from next-token log-probabilities under counterbalanced label mappings and pre-specified health gates. A first-signal pilot on an open-weight instruction-tuned model reveals the central measurement problem. Although all pre-specified health gates passed, the binary-conditioned distributions were near-deterministic for 17 of 18 item pairs under the direct-evaluation framing and 7 of 8 under the persona framing. Label assignment materially changed several mapping-specific QQ verdicts, and no item was certified as residually contextual. Thus, under the tested conditions, the observed QQ outcomes did not uniquely identify a response mechanism in the presence of a saturated and label-sensitive measurement interface. The main implication is methodological: next-token probabilities should not be interpreted as survey-response distributions without first establishing adequate dispersion. We therefore argue that saturation screening and label counterbalancing should precede structural interpretation in distribution-level audits of LLM judgments.

Figures

Figures reproduced from arXiv: 2607.17219 by Pilsung Kang.

Figure 1
Figure 1. Figure 1: The audit pipeline (left column) and the outcomes it produced in this pilot (right column, [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Empirical cumulative distribution function (CDF) of the distance from determinism at [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Label effect under counterbalancing: per-item [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

27 extracted references · 5 canonical work pages

  1. [1]

    David W. Moore. Measuring New Types of Question-Order Effects: Additive and Subtractive. Public Opinion Quarterly, 66(1):80–91, 2002. doi: 10.1086/338631

  2. [2]

    Busemeyer

    Zheng Wang and Jerome R. Busemeyer. A Quantum Question Order Model Supported by Empirical Tests of an A Priori and Precise Prediction.Topics in Cognitive Science, 5(4): 689–710, 2013. doi: 10.1111/tops.12040. URL https://onlinelibrary.wiley.com/doi/abs/ 10.1111/tops.12040

  3. [3]

    Shiffrin, and Jerome R

    Zheng Wang, Tyler Solloway, Richard M. Shiffrin, and Jerome R. Busemeyer. Context Effects Produced by Question Orders Reveal Quantum Nature of Human Judgments.Proceedings of the National Academy of Sciences, 111(26):9431–9436, 2014. doi: 10.1073/pnas.1407756111

  4. [4]

    Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey Design.Transactions of the Association for Computational Linguistics, 12:1011–1026, 2024

    Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkar, and Graham Neubig. Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey Design.Transactions of the Association for Computational Linguistics, 12:1011–1026, 2024. doi: 10.1162/tacl a 00685. URLhttps://aclanthology.org/2024.tacl-1.56/

  5. [5]

    Prompt Perturbations Reveal Human- Like Biases in Large Language Model Survey Responses

    Jens Rupprecht, Georg Ahnert, and Markus Strohmaier. Prompt Perturbations Reveal Human- Like Biases in Large Language Model Survey Responses. InProceedings of the Seventh Workshop on Natural Language Processing and Computational Social Science (NLP+CSS), pages 1–21, San Diego, CA, 2026. Association for Computational Linguistics

  6. [6]

    CAPE: Context-Aware Personality Evaluation Framework for Large Language Models

    Jivnesh Sandhan, Fei Cheng, Tushar Sandhan, and Yugo Murawaki. CAPE: Context-Aware Personality Evaluation Framework for Large Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 10648–10662. Association for Computational Linguistics, 2025

  7. [7]

    World Scientific, Singapore, 2023

    Aleksandr Lebedev and Andrei Khrennikov.Quantum-like Modeling of the Order Effect in Decision Making: POVM Viewpoint on the Wang–Busemeyer QQ-equality, pages 123–128. World Scientific, Singapore, 2023. doi: 10.1142/9789811275999 0010. 21

  8. [8]

    Clopper and Egon S

    Charles J. Clopper and Egon S. Pearson. The Use of Confidence or Fiducial Limits Illustrated in the Case of the Binomial.Biometrika, 26(4):404–413, 1934

  9. [9]

    Qwen3 Technical Report

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 Technical Report. arXiv preprint arXiv:2505.09388, 2025

  10. [10]

    Modeling Combination of Question Order Ef- fect, Response Replicability Effect, and QQ-equality with Quantum Instruments.Jour- nal of Mathematical Psychology, 100:102491, 2021

    Masanao Ozawa and Andrei Khrennikov. Modeling Combination of Question Order Ef- fect, Response Replicability Effect, and QQ-equality with Quantum Instruments.Jour- nal of Mathematical Psychology, 100:102491, 2021. ISSN 0022-2496. doi: https://doi.org/ 10.1016/j.jmp.2020.102491. URL https://www.sciencedirect.com/science/article/pii/ S0022249620301152

  11. [11]

    Dzhafarov, Janne V

    Ehtibar N. Dzhafarov, Janne V. Kujala, and Victor H. Cervantes. Contextuality-by-Default: A Brief Overview of Ideas, Concepts, and Terminology. In Harald Atmanspacher, Thomas Filk, and Emmanuel Pothos, editors,Quantum Interaction, pages 12–23, Cham, 2016. Springer International Publishing

  12. [12]

    Kujala and Ehtibar N

    Janne V. Kujala and Ehtibar N. Dzhafarov. Proof of a Conjecture on Contextuality in Cyclic Systems with Binary Variables.Foundations of Physics, 46(3):282–299, Mar 2016. ISSN 1572-

  13. [13]

    Failure of Contextual Invariance in Large Language Models, 2026

    Sagar Kumar, Ariel Flint, Luca Maria Aiello, and Andrea Baronchelli. Failure of Contextual Invariance in Large Language Models, 2026. URLhttps://arxiv.org/abs/2603.23485

  14. [14]

    Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study, 2026

    Thanh Luong Tuan. Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study, 2026. URLhttps://arxiv.org/abs/2607.11948

  15. [15]

    Pothos, and Andy J

    Pegah Imannezhad, Emmanuel M. Pothos, and Andy J. Wills. Divergent Patterns of Proba- bilistic Reasoning in Humans and GPT-5.Frontiers in Psychology, 17:1782184, 2026. ISSN 1664-1078. doi: 10.3389/fpsyg.2026.1782184. URL https://www.frontiersin.org/journals/ psychology/articles/10.3389/fpsyg.2026.1782184

  16. [16]

    Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Human and Artificial Cognition.Entropy, 28(6):622, June 2026

    Diederik Aerts, Jonito Aerts Argu¨ elles, Lester Beltran, Suzette Geriente, Roberto Leporini, Massimiliano Sassoli de Bianchi, and Sandro Sozzo. Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Human and Artificial Cognition.Entropy, 28(6):622, June 2026. ISSN 1099-4300. doi: 10.3390/e28060622. URL http://dx.doi.org/1...

  17. [17]

    Quantum-like Contextuality in Large Language Models.Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences, 481(2319):20240399, 08 2025

    Kin Ian Lo, Mehrnoosh Sadrzadeh, and Shane Mansfield. Quantum-like Contextuality in Large Language Models.Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences, 481(2319):20240399, 08 2025. ISSN 1364-5021. doi: 10.1098/rspa.2024.0399. URL https://doi.org/10.1098/rspa.2024.0399

  18. [18]

    Large Language Models Are Not Robust Multiple Choice Selectors

    Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. Large Language Models Are Not Robust Multiple Choice Selectors. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=shr9PXz7T0

  19. [19]

    Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions

    Pouya Pezeshkpour and Estevam Hruschka. Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions. InFindings of the Association for Computational Linguistics: NAACL 2024, pages 2006–2017, 2024. 22

  20. [20]

    My Answer is C

    Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul R¨ ottger, Frauke Kreuter, Dirk Hovy, and Barbara Plank. “My Answer is C”: First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models. InFindings of the Association for Com- putational Linguistics: ACL 2024, pages 7407–7416. Association for Computational Linguistics, 2024

  21. [21]

    doi: https://doi.org/10.1002/andp.19504430510

    Gerhart L¨ uders.¨Uber die Zustands¨ anderung durch den Meßprozess.Annalen der Physik, 443 (5-8):322–328, 1950. doi: https://doi.org/10.1002/andp.19504430510

  22. [22]

    Understanding the Effects of RLHF on LLM Generalisation and Diversity

    Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. Understanding the Effects of RLHF on LLM Generalisation and Diversity. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=PXD3FAVHJT

  23. [23]

    Does Writing with Language Models Reduce Content Diversity? InThe Twelfth International Conference on Learning Representations, 2024

    Vishakh Padmakumar and He He. Does Writing with Language Models Reduce Content Diversity? InThe Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Feiz5HtCD0

  24. [24]

    Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

    Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empiri...

  25. [25]

    A" or "B

    Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec,...

  26. [2023]

    doi: 10.18653/v1/2023.emnlp-main.330

    Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.330. URL https://aclanthology.org/2023.emnlp-main.330/

  27. [9516]

    URL https://doi.org/10.1007/s10701-015-9964-8

    doi: 10.1007/s10701-015-9964-8. URL https://doi.org/10.1007/s10701-015-9964-8