Pith. sign in

REVIEW 4 major objections 4 minor 66 references

This paper argues that the answer-logit difference in multimodal language models behaves as a monotonic readout of a latent decision variable, sufficient in principle to compute posterior confidence in simple perceptual and memory-based dec

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 06:30 UTC pith:NGHBN3EM

load-bearing objection A careful, honest empirical paper that imports SDC from neuroscience to ask what answer-logit confidence represents in multimodal LLMs; the folded-X results are real and interesting, but the perceptual-task evidence leans on the untested assumption that nuisance variation is pure internal noise. the 4 major comments →

arxiv 2607.12447 v2 pith:NGHBN3EM submitted 2026-07-14 cs.LG cs.AI

The Computational Basis of Confidence in Large Language Models

classification cs.LG cs.AI
keywords confidenceanswer logitslatent decision variablestatistical decision confidencefolded-X patternmultimodal language modelscalibrationabstention
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks what a language model's confidence signal actually represents, not just whether it predicts correct answers. It claims that the difference between the logits of competing answer tokens (LD) behaves as a monotonic readout of a latent decision variable—a noisy internal representation of task-relevant evidence that, under a normative model, is sufficient in principle to compute the posterior probability that the chosen answer is correct. Across three perceptual discrimination tasks and a memory-based city-population comparison, four qualitative signatures predicted by statistical decision confidence hold, including the diagnostic folded-X pattern. The same signal drives near-optimal abstention: models withhold low-confidence answers and raise committed accuracy to near ceiling. In complex visual reasoning, LD still predicts correctness beyond objective difficulty, but the full folded-X signature is absent, marking the boundary where no explicit normative process model is available.

Core claim

Across three perceptual discrimination tasks and a memory-based city-population comparison, the answer-logit difference (LD) satisfies four qualitative signatures of statistical decision confidence: monotonic psychometric choice, signed LD tracking signed stimulus strength, |LD| predicting correctness at fixed stimulus strength, and the folded-X pattern. This holds in three non-reasoning models and one reasoning model, with a single exception. In the memory task, the folded-X is sharp despite a shallow psychometric. In complex visual reasoning (CLEVR), |LD| still predicts correctness beyond blur and scene controls, but the folded-X is absent; the authors treat this as inconclusive because no

What carries the argument

The central object is the answer-logit difference LD = ℓ_A − ℓ_B, candidate readout of a latent decision variable d = Δ + ε. Statistical decision confidence defines confidence as P(correct|d), so d is a sufficient statistic; any monotonic readout inherits its qualitative signatures. The key experimental device substitutes across-exemplar nuisance variation (same objective strength, different task-irrelevant features) for internal sensory noise, allowing deterministic LLMs to exhibit trial-like fluctuations at fixed strength and making the fixed-strength and folded-X tests possible.

Load-bearing premise

The argument depends on treating across-exemplar nuisance variation in stimuli (same objective strength, different task-irrelevant features) as equivalent to internal sensory noise; if nuisance features change objective difficulty rather than acting as pure noise, the fixed-strength and folded-X tests are confounded.

What would settle it

A concrete falsifier: construct two sets of stimuli at the same nominal stimulus strength—one with nuisance features that are genuinely task-irrelevant and one with nuisance features that objectively change discriminability (for example, altering effective contrast). If the folded-X and fixed-strength effects appear in both sets, or disappear when nuisance features are truly task-irrelevant, the claim that LD reads out a latent decision variable would be undermined. Alternatively, independently varying evidence reliability while holding LD fixed should distinguish Bayesian confidence from raw

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Answer-logit difference can be interpreted as a latent decision-variable readout in perceptual and memory tasks, so confidence in these settings is a normative quantity, not just an empirical predictor.
  • Abstention decisions track LD near-optimally: committing only high-|LD| responses raised accuracy from 92.2% to 99.95%, and the policy overlapped 91.6% with an ideal LD-threshold policy.
  • In complex visual reasoning, |LD| still predicts correctness beyond blur and scene controls, but the absence of the folded-X is inconclusive without a specified normative process model.
  • The framework extends beyond vision to memory-based language tasks, suggesting it applies to any discrete-response LLM task where a normative process model can be specified.
  • The signatures cannot currently distinguish Bayesian confidence from raw evidence strength; that requires independently manipulating evidence reliability and stimulus strength.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If LD truly reads out a latent decision variable, post-hoc calibration is a second-order fix: temperature scaling adjusts the mapping from LD to accuracy but does not create or destroy the sufficiency of LD itself.
  • The framework suggests a direct intervention: design nuisance features that push the internal evidence while leaving objective strength fixed, and measure whether the resulting LD distribution follows the SDC geometry; this would test the equivalence of nuisance variation and internal noise.
  • The CLEVR boundary implies that reasoning tasks may need task-specific normative models; a testable extension is to run the same four-signature analysis on other tasks with a known generative evidence axis (for example, arithmetic problems with controlled difficulty) and compare the presence of the folded-X.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper asks whether answer-logit differences (LD) in multimodal LLMs behave as a readout of a latent decision variable sufficient to compute normative posterior confidence, rather than as a heuristic preference score. Using statistical decision confidence (SDC), the authors test four qualitative signatures — psychometric function, signed LD vs. signed stimulus strength, fixed-strength correctness prediction, and the folded-X correct/error interaction — across three synthetic perceptual tasks (contrast, size, shape) and a memory-based city-population comparison, in four models (Qwen 2.5 7B, Gemma 3 12B, Gemma 4 12B, Gemini Flash 3). They report that LD satisfies all four signatures in perceptual and memory tasks (with one exception: Gemma 3 contrast error branch flat), that abstention behavior in Qwen tracks |LD| near-optimally, and that in CLEVR under blur LD predicts correctness beyond objective difficulty but the folded-X is absent. The authors conclude that LD is a monotonic readout of the latent decision variable prescribed by the experimenter-defined normative model.

Significance. If the interpretation holds, this is a significant step: it moves LLM confidence research beyond calibration and AUROC toward a computational-level account, importing a normative framework from neuroscience and providing a unifying language for biological and artificial confidence. The study has substantial strengths: very large trial counts (e.g., 63,500 contrast trials), tests across four models including a reasoning model, pre-specified qualitative signatures, held-out calibration for temperature scaling, and a behavioral consequence (abstention) that links the internal signal to action. The authors are appropriately cautious about the CLEVR boundary case and about BCH vs. CRES. However, the central claim rests on an untested equivalence between across-exemplar nuisance variation and internal sensory noise; if that equivalence fails, the fixed-strength and folded-X results are confounded by genuine within-bin difficulty variation.

major comments (4)
  1. [Methods (Shape categorization); Table 1A] The load-bearing assumption is that nuisance variation across exemplars plays the same role as internal noise ε in the normative model d = Δ + ε. This is asserted in Figure 1 and Methods, but not established. In the contrast task, independent band-pass noise patches matched for RMS contrast still differ in spectral energy in frequency bands the model is sensitive to, so the true objective difficulty varies within a nominal Δ bin. In the size and shape tasks, positional and side-length jitter change low-level cues. Thus the fixed-strength regression (correct ~ C(|stimulus|,bin) + |LD|, Table 1B) may be significant simply because |LD| tracks these within-bin objective difficulty differences, and the folded-X (Table 1C) may arise generically for any calibrated confidence score when difficulty varies. The paper's Supplemental Discussion concedes that if task-irrelevant features enter the dec
  2. [Discussion (BCH/CRES; sufficiency); Supplemental Discussion] For the shape task, stimulus strength is defined relative to the model's own empirically estimated PSE (m* = 0.73 for Qwen), so the signed-evidence axis is not an experimenter-defined objective axis. Correctness is also defined relative to that internal boundary, making trials near m* ambiguous by construction; in that regime |LD| will naturally correlate with accuracy. The psychometric curve then partly measures the model against its own fitted bias, rather than against a ground-truth boundary. The same issue appears in the model-relative size analysis for Gemma 3 in the Supplemental Results. This circularity weakens the inference that LD reads out the experimenter-defined normative latent variable. The authors should provide analyses using an objective category boundary (e.g., m = 0.5 or a physical axis) or explicitly justify why the PSE-relative axis does not induce the reported signa
  3. [Results (CLEVR); Table 8] The conclusion that LD is 'sufficient in principle to compute the posterior confidence' is stronger than what the four signatures establish. The signatures are qualitative and, as the authors note, cannot distinguish Bayesian confidence from raw evidence strength (CRES). More importantly, monotonicity of LD with a latent variable does not prove sufficiency: a heuristic score that is a monotone function of a decision variable contaminated by task-irrelevant features could satisfy all four signatures while not being sufficient for the experimenter-defined posterior. The recognition-memory thought experiment in the Supplemental Discussion makes this possible mismatch explicit. The manuscript should either weaken the claim to 'behaves as a monotonic readout consistent with SDC under the experimenter-defined model' or provide an additional test, such as showing that LD is independent of nuisa
  4. [Figure 2; Methods (Abstention task)] The CLEVR analyses involve several post-hoc restrictions: Qwen is evaluated on yes-truth trials only because of a strong affirmative bias; the Gemini Flash 3 count difficulty-controlled model is non-estimable due to collinearity; Gemma 4 size task is excluded. These are disclosed, but the generality of the 'adds information beyond controls' claim should be calibrated accordingly. Also, in Table 8 the Gemini Flash 3 count difficulty-controlled coefficient shows a non-significant Wald result (†) while the likelihood-ratio test is reported as significant; the reader should be told whether the LR test is on the same nested model and how collinearity affects it. The absence of the folded-X in CLEVR is correctly interpreted as non-diagnostic, but this shifts the burden to the perceptual/memory results, which are exactly where the nuisance-variation equivalence is most questionable.
minor comments (4)
  1. [Methods (Models; reasoning-cut readout)] Figure 2's calibration curves use a temperature T fit on non-held-out trials; please clarify whether the ECE minimization is on the training set and whether the held-out set is strictly disjoint from any analysis that later conditions on correctness (it appears to be, but the text could be more explicit).
  2. [Table 7 note] The '90% reasoning-cut readout' for Gemini Flash 3 is described briefly; please specify how the truncated trace is turned into a one-letter forced choice and whether the model's answer logits after the cut are obtained with temperature or greedy decoding. A few more implementation details would aid reproducibility.
  3. [General] Table 7 reports N=9,459 for the fixed-bin logistic model but the psychometric N is 16,800; state the exclusion criteria (e.g., removal of bins with too few trials) to avoid apparent discrepancies.
  4. [General] The paper uses 'strength' inconsistently (objective stimulus strength vs. |LD| as decision strength). Consider standardizing terminology, e.g., 'physical strength' and 'decision strength'.

Circularity Check

2 steps flagged

Core SDC signatures are externally derived and not circular; minor self-referential elements in the shape-task PSE and the signed X-pattern illustration keep the score at 2.

specific steps
  1. fitted input called prediction [Methods, Shape categorization; Results, Encoding of Evidence]
    "Stimulus strength was defined relative to each model's empirically estimated angular–round boundary m*, as Δ = m − m*, so that Δ = 0 marks the model's point of subjective equality (m* = 0.73 for Qwen, 0.38 for Gemma 3). ... the model demonstrated orderly psychometric behavior, satisfying the first core signature of evidence encoding."

    For the shape task, the signed stimulus axis and the definition of 'correct' are built from the model's own fitted point of subjective equality. The psychometric crossing at Δ=0 is therefore guaranteed by construction, and 'correctness' is defined relative to the model's subjective boundary rather than an independent ground truth. This makes the first encoding signature for shape partly a restatement of the fitted PSE rather than an independent SDC prediction. It does not force the fixed-strength or folded-X signatures, which are the load-bearing tests.

  2. self definitional [Figure 4 caption, top panel; Methods, City-population task]
    "Top: Full signed X-pattern. The task-specific answer-logit difference LD is plotted against the signed stimulus variable, separately for correct and error trials. Correct trials lie on the stimulus-congruent side of LD=0, whereas error trials lie on the opposite side. ... Accuracy was computed by comparing the sign of LD with the sign of Δ."

    Because correctness is operationally defined as sign(LD) matching sign(Δ), the statement that correct trials lie on the stimulus-congruent side of LD=0 and error trials on the opposite side is true by definition. The signed X-pattern panel is therefore illustrative rather than evidential. The diagnostic folded-X is the separate magnitude regression (|LD| vs |Δ| by correctness), which is not tautological and is the actual test used in Table 1(C).

full rationale

The central SDC derivation is not circular. The four signatures are derived from an external normative model (d=Δ+ε, SDC=P(correct|d)) and are tested against observable answer-logit differences and experimenter-controlled stimulus strengths; the core signature regressions contain no fitted parameters. The fixed-strength and folded-X tests are nontrivial consequences of the normative model, not identities. Self-citations (Kumaran et al. 2026a,b) support motivational or auxiliary claims and are not load-bearing for the derivation. The paper explicitly disclaims uniqueness of the normative model and acknowledges it cannot separate BCH from CRES, so no uniqueness is imported from the authors' prior work. Two minor self-referential elements prevent a score of 0: the shape-task stimulus strength and correctness are defined using the model's own fitted PSE, making the shape psychometric signature partly by construction, and the 'signed X-pattern' panel is definitional given the operational definition of correctness. However, these affect a peripheral illustration and one of three perceptual tasks; the contrast and size tasks, which have objective ground truth and are the main load-bearing evidence, do not reduce to fitted inputs. The abstention threshold and temperature calibration are honestly described and do not affect the central signatures. Overall, no central prediction reduces to a fit or to a self-citation chain; the observed circularity is mild and localized.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The central interpretation rests on the SDC normative model and the equivalence between nuisance-image variation and internal decision noise; these are domain assumptions imported from neuroscience rather than derived for LLMs. The shape task additionally imports a fitted PSE into the definition of correctness. No new physical or computational entities are introduced.

free parameters (3)
  • Calibration temperature T = per task, not listed numerically
    Fitted on non-held-out trials by minimizing ECE; used for calibration displays only, not for the four signature tests.
  • Shape PSE m* = 0.73 (Qwen), 0.38 (Gemma 3)
    Fitted to each model's choice data and used to define signed strength and 'correct' labels in the shape task; re-centering on the model's own boundary weakens the objectivity of that task.
  • Abstention threshold (60%) = 0.60
    Prompt parameter selected empirically to create enough abstention variability; affects abstention results, not the SDC signatures.
axioms (4)
  • domain assumption Normative model d = Δ + ε, with ε internal noise, and a monotonic f such that LD = f(d).
    This is the SDC framework borrowed from neuroscience; the paper tests its qualitative signatures rather than proving them for LLMs (Methods, Normative model).
  • ad hoc to paper Across-exemplar nuisance variation plays the role of internal noise ε.
    Deterministic LLMs have no trial-to-trial noise; the paper substitutes variation in nuisance image features. If this variation changes objective difficulty rather than acting as pure internal noise, the fixed-strength and folded-X tests are confounded (Figure 1 caption; Methods).
  • domain assumption Additive-noise decision process with boundary-defined response categories yields the folded-X geometry.
    The folded-X is a prediction of a specific process model; the paper cites Adler & Ma for robustness in this regime but does not test alternative noise distributions (Discussion).
  • ad hoc to paper For shape, correctness is defined relative to the model's own PSE m*, not an objective category boundary.
    Because the morph continuum has no ground-truth boundary, the paper fits m* and labels trials correct by sign(m - m*), making 'accuracy' model-relative (Methods, Shape categorization).

pith-pipeline@v1.3.0-alltime-deepseek · 20971 in / 14244 out tokens · 146506 ms · 2026-08-02T06:30:16.357507+00:00 · methodology

0 comments
read the original abstract

Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is calibrated, leaving open a more fundamental question: what does the confidence signal itself represent? Answer logits may reflect a latent decision variable sufficient to compute normative confidence, or instead a heuristic preference signal that combines the available evidence in a non-Bayesian manner. We address this using statistical decision confidence (SDC), a normative framework from computational neuroscience. Treating the answer-logit difference (LD) as a candidate readout of the latent decision variable, we test the qualitative signatures predicted by SDC. Across three perceptual discrimination tasks and a memory-based decision task, spanning three multimodal non-reasoning models and one reasoning model, LD satisfied these signatures -- including the diagnostic correct/error folded-X pattern -- showing that, in these settings, answer logits behave as monotonic readouts of a latent decision variable rather than heuristic preference scores. In complex visual reasoning, LD continued to predict correctness beyond objective task difficulty, but the full geometric signatures of SDC were absent, illustrating the current boundary of the framework when explicit normative process models are unavailable. These results provide a computational account of confidence in multimodal language models, delineate when answer logits behave as readouts of a latent decision variable, and establish SDC as a unifying framework for studying confidence across biological and artificial intelligence.

Figures

Figures reproduced from arXiv: 2607.12447 by Dharshan Kumaran, Maks Ovsjanikov, Nathaniel Daw, Petar Veli\v{c}kovi\'c, Viorica Patraucean.

Figure 1
Figure 1. Figure 1: Schematic of the normative perceptual decision model used to define statistical decision [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Qwen 2.5 7B: Choice, calibration, and evidence encoding across visual decision tasks. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Qwen 2.5 7B: Answer-logit confidence predicts correctness within fixed stimulus-strength bins. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Qwen 2.5 7B: Decision-strength and signed evidence structure in answer logits. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Qwen 2.5 7B, population QA: answer-logit confidence tracks correctness despite a shallow [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: CLEVR count and existence tasks under parametric gaussian blur. [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Example stimuli from simple perceptual decision making tasks. [PITH_FULL_IMAGE:figures/full_fig_p023_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Gemma 3 12B: Choice behaviour, LD-stimulus strength plot and calibration across visual [PITH_FULL_IMAGE:figures/full_fig_p023_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Answer-logit confidence predicts correctness within fixed stimulus-strength bins in Gemma [PITH_FULL_IMAGE:figures/full_fig_p024_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Gemma 3 12B: Decision-strength and signed evidence structure in answer logits. [PITH_FULL_IMAGE:figures/full_fig_p025_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Gemma 3 12B population QA: Results Top left: psychometric curve for population comparison trials. The signed stimulus is the log2 population ratio, log2 (A/B), so positive values favour answer A. Top right: signed LD is correlated with binned stimulus strength. Bottom left: fixed-evidence analysis; trials were grouped by absolute population evidence and then binned by raw logit-difference magnitude. Botto… view at source ↗
Figure 12
Figure 12. Figure 12: Perceptual tasks and Gemma 4 12B Top: psychometric curves for contrast discrimination and shape categorization. Points show binned choice probabilities with SEM error bars; solid lines show fitted logistic psychometric curves and dotted vertical lines mark fitted thresholds. Middle: fixed stimulus-strength analyses. Trials were grouped by objective stimulus strength and then binned by answer-logit confide… view at source ↗
Figure 13
Figure 13. Figure 13: Contrast Task and Gemini Flash 3 – a reasoning model [PITH_FULL_IMAGE:figures/full_fig_p028_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Calibrated confidence in CLEVR for Qwen 2.5 7B and Gemini Flash 3 [PITH_FULL_IMAGE:figures/full_fig_p029_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Answer-logit confidence during the perceptual contrast task (LD) predicts abstention behavior in Qwen 2.5 7B 29 [PITH_FULL_IMAGE:figures/full_fig_p029_15.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

66 extracted references · 21 linked inside Pith

  1. [1]

    Nature , volume=

    Frontal cortex neuron types categorically encode single decision variables , author=. Nature , volume=. 2019 , publisher=

  2. [2]

    Neuron , volume=

    Signatures of a statistical computation in the human sense of confidence , author=. Neuron , volume=. 2016 , publisher=

  3. [3]

    Neural Computation , volume=

    A mathematical framework for statistical decision confidence , author=. Neural Computation , volume=. 2016 , publisher=

  4. [4]

    , author=

    The physics of optimal decision making: a formal analysis of models of performance in two-alternative forced-choice tasks. , author=. Psychological review , volume=. 2006 , publisher=

  5. [5]

    PLoS computational biology , volume=

    The folded X-pattern is not necessarily a statistical signature of decision confidence , author=. PLoS computational biology , volume=. 2019 , publisher=

  6. [6]

    Lawrence and Girshick, Ross , title =

    Johnson, Justin and Hariharan, Bharath and van der Maaten, Laurens and Fei-Fei, Li and Zitnick, C. Lawrence and Girshick, Ross , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =. 2017 , doi =

  7. [7]

    Kumaran, Dharshan and Conmy, Arthur and Barbero, Federico and Osindero, Simon and Patraucean, Viorica and Veli. How do. arXiv preprint arXiv:2603.17839 , year=

  8. [8]

    arXiv preprint arXiv:2304.13734 , year=

    The internal state of an LLM knows when it's lying , author=. arXiv preprint arXiv:2304.13734 , year=

  9. [9]

    arXiv preprint arXiv:2503.19786 , year=

    Gemma 3 technical report , author=. arXiv preprint arXiv:2503.19786 , year=

  10. [10]

    Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

    A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference , author=. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

  11. [11]

    Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

    The validation gap: A mechanistic analysis of how language models compute arithmetic but fail to validate it , author=. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

  12. [12]

    Advances in Neural Information Processing Systems , volume=

    Truth is universal: Robust detection of lies in llms , author=. Advances in Neural Information Processing Systems , volume=

  13. [13]

    arXiv preprint arXiv:2212.03827 , year=

    Discovering latent knowledge in language models without supervision , author=. arXiv preprint arXiv:2212.03827 , year=

  14. [14]

    arXiv preprint arXiv:2501.19306 , year=

    Sets: Leveraging self-verification and self-correction for improved test-time scaling , author=. arXiv preprint arXiv:2501.19306 , year=

  15. [15]

    , author=

    Self-evaluation of decision-making: A general Bayesian framework for metacognitive computation. , author=. Psychological review , volume=. 2017 , publisher=

  16. [16]

    arXiv preprint arXiv:2311.08298 , year=

    A survey of language model confidence estimation and calibration , author=. arXiv preprint arXiv:2311.08298 , year=

  17. [17]

    arXiv preprint arXiv:2404.15255 , year=

    How to use and interpret activation patching , author=. arXiv preprint arXiv:2404.15255 , year=

  18. [18]

    International Conference on Learning Representations , year=

    Large Language Models Cannot Self-Correct Reasoning Yet , author=. International Conference on Learning Representations , year=

  19. [19]

    arXiv preprint arXiv:1705.03551 , year=

    Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension , author=. arXiv preprint arXiv:1705.03551 , year=

  20. [20]

    arXiv preprint arXiv:2207.05221 , year=

    Language models (mostly) know what they know , author=. arXiv preprint arXiv:2207.05221 , year=

  21. [21]

    Transactions of the Association for Computational Linguistics , volume=

    When can llms actually correct their own mistakes? a critical survey of self-correction of llms , author=. Transactions of the Association for Computational Linguistics , volume=

  22. [22]

    Nature , volume=

    Neural correlates, computation and behavioural impact of decision confidence , author=. Nature , volume=. 2008 , publisher=

  23. [23]

    Philosophical Transactions of the Royal Society B: Biological Sciences , volume=

    A computational framework for the study of confidence in humans and animals , author=. Philosophical Transactions of the Royal Society B: Biological Sciences , volume=. 2012 , publisher=

  24. [24]

    science , volume=

    Representation of confidence associated with a decision by neurons in the parietal cortex , author=. science , volume=. 2009 , publisher=

  25. [25]

    Advances in Neural Information Processing Systems , volume=

    Inference-time intervention: Eliciting truthful answers from a language model , author=. Advances in Neural Information Processing Systems , volume=

  26. [26]

    arXiv preprint arXiv:2402.12563 , year=

    Confidence matters: Revisiting intrinsic self-correction capabilities of large language models , author=. arXiv preprint arXiv:2402.12563 , year=

  27. [27]

    arXiv preprint arXiv:2406.15673 , year=

    Large language models have intrinsic self-correction ability , author=. arXiv preprint arXiv:2406.15673 , year=

  28. [28]

    arXiv preprint arXiv:2510.04013 , year=

    LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization , author=. arXiv preprint arXiv:2510.04013 , year=

  29. [29]

    Advances in neural information processing systems , volume=

    Self-refine: Iterative refinement with self-feedback , author=. Advances in neural information processing systems , volume=

  30. [30]

    Advances in neural information processing systems , volume=

    Locating and editing factual associations in gpt , author=. Advances in neural information processing systems , volume=

  31. [31]

    arXiv preprint arXiv:2410.02707 , year=

    Llms know more than they show: On the intrinsic representation of llm hallucinations , author=. arXiv preprint arXiv:2410.02707 , year=

  32. [32]

    Nature neuroscience , volume=

    Confidence and certainty: distinct probabilistic quantities for different goals , author=. Nature neuroscience , volume=. 2016 , publisher=

  33. [33]

    Nature , volume=

    Error correction time without external error signals , author=. Nature , volume=. 1966 , doi=

  34. [34]

    Nature Machine Intelligence , volume=

    What Large Language Models Know and What People Think They Know , author=. Nature Machine Intelligence , volume=

  35. [35]

    arXiv preprint arXiv:2305.14975 , year=

    Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback , author=. arXiv preprint arXiv:2305.14975 , year=

  36. [36]

    ICLR , year=

    Interpretability in the wild: a circuit for indirect object identification in gpt-2 small , author=. ICLR , year=

  37. [37]

    Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

    Large language models are better reasoners with self-verification , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

  38. [38]

    arXiv preprint arXiv:2306.13063 , year=

    Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms , author=. arXiv preprint arXiv:2306.13063 , year=

  39. [39]

    Psychological Review , volume=

    The neural basis of error detection: Conflict monitoring and the error-related negativity , author=. Psychological Review , volume=. 2004 , doi=

  40. [40]

    arXiv preprint arXiv:2505.14489 , year=

    Reasoning Models Better Express Their Confidence , author=. arXiv preprint arXiv:2505.14489 , year=

  41. [41]

    arXiv preprint arXiv:2309.16042 , year=

    Towards best practices of activation patching in language models: Metrics and methods , author=. arXiv preprint arXiv:2309.16042 , year=

  42. [42]

    arXiv preprint arXiv:2510.20487 , year=

    Steering Evaluation-Aware Language Models To Act Like They Are Deployed , author=. arXiv preprint arXiv:2510.20487 , year=

  43. [43]

    arXiv preprint arXiv:2312.06681 , year=

    Steering llama 2 via contrastive activation addition , author=. arXiv preprint arXiv:2312.06681 , year=

  44. [44]

    arXiv preprint arXiv:2410.12877 , year=

    Improving instruction-following in language models through activation steering , author=. arXiv preprint arXiv:2410.12877 , year=

  45. [45]

    arXiv preprint arXiv:2308.10248 , year=

    Steering language models with activation engineering , author=. arXiv preprint arXiv:2308.10248 , year=

  46. [46]

    arXiv preprint arXiv:2510.07364 , year=

    Base Models Know How to Reason, Thinking Models Learn When , author=. arXiv preprint arXiv:2510.07364 , year=

  47. [47]

    Know When You're Wrong: Aligning Confidence with Correctness for

    Xie, Xiaohu and Liu, Xiaohu and Yao, Benjamin , journal=. Know When You're Wrong: Aligning Confidence with Correctness for

  48. [48]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Confidence vs critique: A decomposition of self-correction capability for llms , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  49. [49]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Understanding the dark side of llms' intrinsic self-correction , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  50. [50]

    Nature Reviews Neuroscience , volume=

    Recognition memory: what are the roles of the perirhinal cortex and hippocampus? , author=. Nature Reviews Neuroscience , volume=. 2001 , publisher=

  51. [51]

    arXiv preprint arXiv:2510.27328 , year=

    A Unified Representation Underlying the Judgment of Large Language Models , author=. arXiv preprint arXiv:2510.27328 , year=

  52. [52]

    arXiv:2507.12638 , year=

    Ward, Jake and Lin, Chuqiao and Venhoff, Constantin and Nanda, Neel , title=. arXiv:2507.12638 , year=

  53. [53]

    , title=

    Gandhi, Kanishk and Chakravarthy, Ayush and Singh, Anikait and Lile, Nathan and Goodman, Noah D. , title=. arXiv:2503.01307 , year=

  54. [54]

    arXiv:2502.04404 , year=

    Yang, Xiao-Wen and Zhu, Xiao-Yu and Wei, Wei-Da and Zhang, De-Chuan and Shao, Jian-Jun and Zhou, Zhi and Guo, Lan-Zhe and Li, Yu-Feng , title=. arXiv:2502.04404 , year=

  55. [55]

    Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

    When not to trust language models: Investigating effectiveness of parametric and non-parametric memories , author=. Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

  56. [56]

    Current Directions in Psychological Science , pages=

    Metacognition and uncertainty communication in humans and large language models , author=. Current Directions in Psychological Science , pages=. 2025 , publisher=

  57. [57]

    Neuron , volume=

    Orbitofrontal cortex is required for optimal waiting based on decision confidence , author=. Neuron , volume=. 2014 , publisher=

  58. [58]

    Cell , volume=

    Behavior-and modality-general representation of confidence in orbitofrontal cortex , author=. Cell , volume=. 2020 , publisher=

  59. [59]

    Causal evidence that language models use confidence to drive behavior , journal =

    Kumaran, Dharshan and Daw, Nathaniel and Osindero, Simon and Veli. Causal evidence that language models use confidence to drive behavior , journal =. 2026 , note =

  60. [60]

    How do LLMs compute verbal confidence? , booktitle =

    Kumaran, Dharshan and Conmy, Arthur and Barbero, Federico and Osindero, Simon and Patraucean, Viorica and Veli. How do LLMs compute verbal confidence? , booktitle =. 2026 , note =

  61. [61]

    Neural computation , volume=

    Limitations of proposed signatures of Bayesian confidence , author=. Neural computation , volume=. 2018 , publisher=

  62. [62]

    , author=

    Modeling perceptual confidence and the confidence forced-choice paradigm. , author=. Psychological review , volume=. 2022 , publisher=

  63. [63]

    Annual Review of Vision Science , volume=

    Visual confidence , author=. Annual Review of Vision Science , volume=. 2016 , publisher=

  64. [64]

    International conference on machine learning , pages=

    On calibration of modern neural networks , author=. International conference on machine learning , pages=. 2017 , organization=

  65. [65]

    2026 , journal =

    Kumaran, Dharshan , title =. 2026 , journal =

  66. [66]

    Challenging the

    Xue, Kai and Shekhar, Medha and Rahnev, Dobromir , journal=. Challenging the. 2024 , doi=