Pith. sign in

REVIEW 4 major objections 4 minor 9 cited by

The paper argues that an LLM's confidence is a separate, measurable capacity from its knowledge.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 17:25 UTC pith:ZYKTUYB5

load-bearing objection The v3 abstract and the v1 body are different papers: the abstract's new metric never appears in the body, and the body's central finding is explicitly retracted by the abstract. the 4 major comments →

arxiv 2603.25112 v3 pith:ZYKTUYB5 submitted 2026-03-26 cs.CL cs.AI

Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory

classification cs.CL cs.AI
keywords LLM confidencemetacognitive efficiencymeta-d' / M-ratioType-2 signal detection theorycalibrationselective predictiontoken log-probabilityfactual question answering
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Current confidence metrics conflate two capacities: whether a model produces correct answers and whether its confidence signal marks the answers it gets wrong. The paper applies Type-2 signal detection theory to token-level log-probabilities, measuring metacognitive efficiency as the ratio of meta-d' to d' across four open-weight models and 224,000 factual QA trials. It claims efficiency is independent of accuracy—the model with the best discrimination has the least usable confidence—varies by knowledge domain, and is dissociated from temperature, which shifts confidence placement while leaving the information content near-flat for three of four models. A revised abstract makes a further claim the printed body does not support: that after correcting a grading bias against human adjudications, the inverse accuracy-efficiency coupling disappears and efficiency should be measured by a model-free information measure. If the core claim holds, model selection for tasks that rely on confidence should use a monitoring metric rather than calibration error.

Core claim

The central claim is that confidence quality and knowledge are separable: a model can be a strong discriminator and a weak monitor, or the reverse. In the body's data, the model with the highest Type-1 sensitivity has the lowest metacognitive efficiency, while a lower-sensitivity model sits near optimal, so two confidence metrics rank the four models in opposite order. Efficiency also differs across knowledge domains—different models are weakest in different subjects—and temperature changes the criterion for reporting confidence without changing how much information confidence contains about correctness for two of the instruction-tuned models. The revised abstract states that these strong em

What carries the argument

Type-2 Signal Detection Theory with meta-d'/d' (M-ratio): token-level normalised log-probability is treated as a graded confidence variable that must discriminate the model's own correct from incorrect answers, and meta-d' is the ideal-observer sensitivity that would produce the observed confidence-by-accuracy contingency table. Dividing by Type-1 d' separates monitoring quality from raw knowledge. The revised abstract replaces this ratio with a model-free information measure (meta-I_2r) for the same purpose, on the ground that open-ended QA lacks the two-alternative decision classical meta-d' assumes; the body defines neither the measure nor the resulting numbers.

Load-bearing premise

The load-bearing premise is that the automated correct/incorrect labels are unbiased; the revised abstract admits that a differential length bias in those labels did not survive human adjudication, and if the labels are wrong the body's efficiency comparisons collapse.

What would settle it

Relabel every trial with the 1,830 human adjudications, recompute metacognitive efficiency by model and domain, and check whether the inverse accuracy-efficiency coupling and the reported z-ROC slope range (0.78–1.18) exist; if the coupling survives, the revised abstract's main correction is false, and if it vanishes, the body's headline results are scorer artifacts.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Evaluation reports should list a metacognitive efficiency metric alongside calibration error, because two models with similar accuracy can have opposite usable-confidence profiles.
  • For selective prediction systems, model choice flips depending on whether confidence is scored by its ranking quality or by metacognitive efficiency; the paper's efficiency-based choice outperforms the ranking-based choice at equal coverage.
  • Domain-level efficiency can hide in aggregate numbers: a model can be metacognitively blind in one subject while efficient in others, so confidence-based deployment should be evaluated per domain.
  • Temperature tuning shifts confidence policy, not capacity, for instruction-tuned models; raising temperature cannot fix a model whose confidence carries little information about correctness.
  • The revised abstract's admission that the inverse accuracy-efficiency coupling does not survive human relabelling implies the printed conclusions are conditional on the original automated scorer.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the proposed information measure replaces M-ratio, the framework could extend beyond two-alternative QA to open-ended generation, where classical meta-d' is not defined.
  • The admitted grader bias is a warning for any LLM-confidence study that relies on automated exact-match scoring: efficiency results should be validated on a human-adjudicated subset before ranking models.
  • The revised abstract's near-perfect rank correlation between metacognitive information and the accuracy gain from abstention, if reproducible, would let practitioners predict coverage-accuracy trade-offs without running the system.
  • Closing the gap between the revised abstract and the printed body is a prerequisite for testing the strongest claims; the z-ROC slopes and cross-model spread cited in the abstract cannot currently be verified from the published analyses.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript (v1 body, v3 abstract) applies Type-2 Signal Detection Theory to measure metacognitive efficiency in four LLMs across 224,000 factual QA trials. The body defines the meta-d'/d' ratio (M-ratio, Eq. 1) and reports that M-ratio varies across models, is domain-specific, and is dissociated from temperature effects; it also reports an inversion between AUROC2 and M-ratio rankings. The v3 abstract, however, introduces a different metric (normalized metacognitive information, meta-I_2r) and five findings (a 1.98-fold range, rank correlations of -0.80 and +0.00, z-ROC slopes 0.78-1.18, Science & Technology as the weakest domain for every model, near-flat metacognitive information under temperature, and rho = +1.00 with abstention gains) that appear nowhere in the body. The abstract also states that M-ratio 'is not well defined for open-ended QA' and that the inverse accuracy-efficiency coupling reported in v1/v2 does not survive corrected labelling. The submitted text is internally inconsistent, and the central claim is not checkable from the analyses presented.

Significance. A valid model-free metric of usable confidence for open-ended QA would be a significant contribution: it could replace calibration metrics that conflate accuracy and confidence informativeness. The v3 abstract's claims, if supported, would be important. The pre-registration, public code/data, and use of permutation nulls and bootstrap CIs are strengths that partially offset concerns about reproducibility. However, the body contains none of the v3 abstract's analyses: meta-I_2r is not defined, z-ROC slopes are not reported, and the abstract's own admission that M-ratio is not well defined for open-ended QA invalidates the body's central metric. The admitted differential length bias in the automated scorer and the statement that the headline finding does not survive relabelling further undermine the empirical contributions. As submitted, the manuscript cannot be evaluated as a coherent paper.

major comments (4)
  1. [Abstract vs. body (§2.2, §4)] The v3 abstract and the v1 body describe different studies. The abstract introduces meta-I_2r and reports findings (factor 1.98, rank correlations -0.80/+0.00, z-ROC slopes 0.78-1.18, Science & Technology as weakest domain for every model, near-flat temperature effect, rho = +1.00 with abstention) that appear nowhere in the body. The body defines only M-ratio (Eq. 1) and reports Tables 2, 6, 7; no meta-I_2r is defined or estimated, and no z-ROC analysis is presented. Moreover, the abstract states M-ratio 'is not well defined for open-ended QA, which lacks a two-alternative Type-1 decision', directly contradicting the body's use of M-ratio as its central metric. The referenced 'version note on page 1' is absent. Without the meta-I_2r analysis, the paper's central claim cannot be checked.
  2. [Abstract vs. §3.1 correctness scoring] The abstract admits that the automated correctness scorer had a 'differential length bias' and that, once corrected against 1,830 human adjudications, 'the inverse accuracy-efficiency coupling reported in v1 and v2 does not survive relabelling'. The body's results (Tables 2-7, Figures 1-3) are all based on the uncorrected labels using exact match plus difflib.SequenceMatcher >= 0.85 (§3.1). No corrected-label analysis is included anywhere in the body. Thus the paper's own headline finding is acknowledged to be an artifact of the grader, and the submitted body still relies on that grader. This is a load-bearing, unresolved problem.
  3. [§4.3, Table 3] The claim that AUROC2 and M-ratio rankings are 'fully inverted' is substantially forced by construction and overstated. M-ratio = meta-d'/d' normalizes out Type-1 sensitivity by division, while AUROC2 covaries with d', so a negative association between the two rankings is mathematically expected when d' varies across models. In addition, Table 3 itself shows Llama-3-Base and Gemma tied at M-ratio 1.048, so the ranking is not fully inverted between ranks 2 and 3. The interpretation that these metrics answer different questions is acceptable, but presenting the inversion as an empirical discovery overstates what the normalization already implies.
  4. [§4.2-§4.6 vs. §5-§6] Reported inferential support for the core hypotheses is weak: H1 is partially supported (one of four models has a CI entirely below 1.0, §4.2), H2 is not met (the pre-registered criterion fails, §4.4), H3 is supported for only two of four models (§4.5), and H4 rests on a single significant pairwise comparison (§4.6). The Discussion and Conclusion nevertheless state these as established findings, including domain-specificity and temperature dissociation. The conclusions should be calibrated to the statistics actually reported; the current text overstates the strength of evidence.
minor comments (4)
  1. [Abstract] The version note on page 1 is referenced but absent; if the version history is important, it should be included or the reference removed.
  2. [§4.3, Table 3] Use of 'fully inverted' is misleading given the tie between Llama-3-Base and Gemma. Consider describing the ranking difference more precisely.
  3. [§5.4] The reference 'Goertz et al. (2024)' is mentioned in the text but not listed in the References section. Please add the full citation.
  4. [§5.1, footnote 1] The footnote reports selective-prediction accuracy at 50% coverage (Gemma 77.4% vs. Mistral 70.7%) but gives no confidence intervals or method details; this is a key practical claim and should be reported with the same inferential rigor as the rest of the paper.

Circularity Check

0 steps flagged

No circular step can be exhibited; abstract/body mismatch is a checkability failure, not a circular derivation.

full rationale

The paper's core body-level quantities are empirical. Equation (1) defines M-ratio as meta-d'/d'; the AUROC2-vs-M-ratio inversion in Table 3 is a data-dependent ranking, and the paper itself explains that AUROC2 inherits Type-1 sensitivity (Section 4.3) rather than pretending the inversion is a derived consequence. Self-citations to Cacioli (2026) are used for prior data collection and Type-1 d' validation, but they do not smuggle in the paper's new claims, and no uniqueness theorem is imported. The v3 abstract's meta-I_2r, z-ROC slopes, and abstention-gain correlation are not defined or derived in the body, and the abstract's own admissions about scorer bias and the ill-definedness of M-ratio for open-ended QA undermine the v1 results; however absence of an implementation or derivation is a correctness/reproducibility problem, not a circularity that can be exhibited as an equality between the claim and its inputs. Under the hard rule requiring a specific reduction, no circular step is identifiable.

Axiom & Free-Parameter Ledger

7 free parameters · 5 axioms · 1 invented entities

Everything downstream depends on the correctness labels and on treating NLP as a confidence variable; the paper's v3 abstract undermines both (scorer bias; M-ratio ill-defined). The v3 metric meta-I_2r is an undefined placeholder in this text.

free parameters (7)
  • Number of confidence bins K = 4 (robustness: 3, 6)
    NLP confidence is discretized into 2×K rating categories; K=4 chosen by hand (§3.2).
  • Bin-edge quantiles = {12.5, 25, 37.5, 50, 62.5, 75, 87.5}th percentiles of NLP at T=1.0 per model×dataset
    Bin edges are fit to each model's own NLP distribution and held across temperatures (§3.2); this data-derived discretization is part of the estimator input.
  • Correctness threshold (difflib ratio) = 0.85
    Answers failing exact-match fall back to SequenceMatcher ratio ≥ 0.85 (§3.1); the threshold is hand-chosen and the scorer is later admitted to carry a differential length bias (v3 abstract).
  • Log-linear correction = +0.5 to all cells (Hautus 1995)
    Applied to avoid degenerate cells in the rating table (§3.2); a standard but non-unique correction.
  • Equal-variance constraint s = 1
    meta-d' MLE assumes equal-variance SDT (§3.2); the paper notes Mistral deviates strongly (s=0.57) yet uses equal-variance meta-d' for it.
  • TOST equivalence bound δ = 0.3
    Temperature-stability hypothesis H3 uses TOST with δ=0.3 (§3.3); chosen by hand.
  • M-ratio instability exclusion = |M-ratio| > 10 excluded
    Bootstrap estimates with |M-ratio|>10 are dropped as unstable (§3.2); post-hoc exclusion.
axioms (5)
  • domain assumption NLP is a monotonically valid graded evidence variable for correctness (functional operationalization of confidence)
    Verified empirically via monotonicity of accuracy across NLP quartiles (§4.1); the metacognitive interpretation is assigned, not measured (§2.3).
  • domain assumption An equal-variance Gaussian SDT generative model underlies Type-1 and Type-2 responses
    Required for d' and meta-d' MLE (metadpy, s=1; §3.2); the paper notes Mistral deviates strongly (s=0.57).
  • domain assumption Meta-d'/M-ratio is well defined for single-response open-ended QA
    The entire body analysis presupposes this; the v3 abstract states it 'is not well defined for open-ended QA, which lacks a two-alternative Type-1 decision' — the body never resolves this.
  • domain assumption Correctness labels from exact match + SequenceMatcher are valid
    Dependent variable for all estimates (§3.1); the v3 abstract reports a differential length bias in this scorer whose correction reversed the headline findings.
  • standard math Bootstrap resampling at trial level accounts for the data structure
    Trial-level resampling with 10,000 resamples (seed 42; §3.2) treats trials as i.i.d.; no clustering by question or model pass is modeled.
invented entities (1)
  • Normalized metacognitive information (meta-I_2r) no independent evidence
    purpose: Claims to quantify Type-2 metacognitive efficiency without requiring a two-alternative Type-1 decision, replacing meta-d'/M-ratio (abstract).
    Referenced only in the abstract; no definition, formula, estimator, or results are present in the body, so it has no falsifiable handle within this text.

pith-pipeline@v1.3.0-alltime-deepseek · 10267 in / 23092 out tokens · 219923 ms · 2026-08-02T17:25:50.869910+00:00 · methodology

0 comments
read the original abstract

Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate two capacities: how much a model knows (Type-1 accuracy) and how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity). We apply Signal Detection Theory to decompose them, treating token-level normalised log-probability as a graded confidence variable and answer correctness as the state to be discriminated. We characterise the Type-2 ROC of this signal, including its unequal-variance structure via z-ROC analysis, and -- because the meta-d' efficiency ratio is not well defined for open-ended QA, which lacks a two-alternative Type-1 decision -- quantify efficiency with a model-free information measure, normalised metacognitive information (meta-I_2r). Across four LLMs and 224,000 factual QA trials we find: (1) metacognitive information varies by a factor of 1.98 across models and is not predicted by accuracy, the rank correlation being -0.80 on TriviaQA and +0.00 on Natural Questions; (2) the confidence signal has model-specific unequal-variance structure (z-ROC slopes 0.78 to 1.18) invisible to calibration metrics, the slope ordering replicating on NQ; (3) efficiency is domain-specific, weakest in Science & Technology for every model; (4) temperature dissociates accuracy from metacognitive information, which stays near-flat for three of four models while accuracy falls; and (5) metacognitive information tracks the accuracy gain from confidence-based abstention exactly (rho = +1.00) while accuracy does not. All estimates carry permutation nulls and bootstrap confidence intervals. This version (v3) corrects a differential length bias in the automated correctness scorer, validated against 1,830 human adjudications; the inverse accuracy-efficiency coupling reported in v1 and v2 does not survive relabelling. See the version note on page 1.

Figures

Figures reproduced from arXiv: 2603.25112 by Jon-Paul Cacioli.

Figure 1
Figure 1. Figure 1: Metacognitive efficiency space. Each point represents a model at [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Domain-specific M-ratio at T=1.0 on TriviaQA. Dashed line: optimal (M=1). Different models exhibit different weakest domains, a pattern invisible to aggregate metrics [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: d ′ and meta-d ′ as a function of temperature on TriviaQA. For Mistral and Gemma, d ′ varies with temperature while meta-d ′ remains stable, dissociating confidence policy from metacognitive capacity. 1,248), indicating limited power rather than absence of effect. Hierarchical Bayesian estimation (HMeta-d; Fleming 2017) may provide the sensitivity needed to confirm domain-specific patterns. 4.5 H3: Tempera… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Do Coding Agents Understand Least-Privilege Authorization?

    cs.CR 2026-05 unverdicted novelty 7.0

    Coding agents struggle to infer least-privilege file permissions by omitting needed accesses while granting unused or sensitive ones, but Sufficiency-Tightness Decomposition improves sensitive-task success by up to 15...

  2. Before You Interpret the Profile: Validity Scaling for LLM Metacognitive Self-Report

    cs.CL 2026-04 conditional novelty 7.0

    Validity indices adapted from clinical assessment classify four frontier LLMs as construct-level invalid on metacognitive probes, with valid models showing positive item-sensitive confidence (r=.18) while invalid ones...

  3. Hypothesis generation and updating in large language models

    cs.LG 2026-05 unverdicted novelty 6.0

    LLMs exhibit Bayesian-like hypothesis updating with strong-sampling bias and an evaluation-generation gap but generalize poorly outside observed data.

  4. Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen

    cs.CL 2026-04 conditional novelty 6.0

    Seven 3-9B instruction-tuned LLMs produce verbal confidence that saturates at high values and fails psychometric validity criteria for Type-2 discrimination under minimal elicitation.

  5. MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition

    cs.AI 2026-04 unverdicted novelty 6.0

    MEDLEY-BENCH reveals an evaluation/control dissociation in AI metacognition where scale improves reflective scoring but not proportional belief revision, with a consistent knowing/doing gap across 35 models.

  6. K-Way Energy Probes for Metacognition Reduce to Softmax in Discriminative Predictive Coding Networks

    cs.LG 2026-04 unverdicted novelty 6.0

    K-way energy probes in discriminative PCNs reduce to a monotone function of the log-softmax margin plus an untrained residual and empirically track below softmax on CIFAR-10.

  7. Disambiguating electrical detection of magnetization dynamics in magnetic insulators

    cond-mat.mes-hall 2026-04 unverdicted novelty 5.0

    Spin pumping and ST-FMR compete with opposite voltage signs in Pt/magnetic-insulator devices; thickness excitation profile and damping, not chirality alone, set which wins.

  8. Disambiguating electrical detection of magnetization dynamics in magnetic insulators

    cond-mat.mes-hall 2026-04 unverdicted novelty 5.0

    Spin pumping and ST-FMR contributions to electrical signals in Pt/magnetic-insulator devices can be separated by geometry and field direction, showing that voltage sign is not a unique indicator of magnon chirality.

  9. Quantisation Reshapes the Metacognitive Geometry of Language Models

    cs.CL 2026-04 unverdicted novelty 5.0

    Quantization restructures domain-level M-ratio metacognitive profiles in LLMs while leaving Type-2 AUROC profiles unchanged.

Reference graph

Works this paper leans on

7 extracted references · 3 linked inside Pith · cited by 8 Pith papers

  1. [3]

    Code and data:https://anonymous.4open.science/r/sdt_calibration

    2Pre-registration:https://osf.io/5q7mt/overview?view_only=bd718de95b6c44ff9c14c1ac424227ba. Code and data:https://anonymous.4open.science/r/sdt_calibration. 9 Stephen M. Fleming. HMeta-d: Hierarchical Bayesian estimation of metacognitive efficiency from confidence ratings.Neuroscience of Consciousness, 2017(1):nix007,

  2. [1995]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chap- lot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7B.arXiv preprint arXiv:...

  3. [1996]

    Can LLMs ex- press their uncertainty? an empirical evaluation of confidence elicitation in LLMs.arXiv preprint arXiv:2306.13063,

    Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. Can LLMs ex- press their uncertainty? an empirical evaluation of confidence elicitation in LLMs.arXiv preprint arXiv:2306.13063,

  4. [2003]

    Gemma 2: Improving open language models at a practical size.arXiv preprint arXiv:2408.00118,

    Gemma Team. Gemma 2: Improving open language models at a practical size.arXiv preprint arXiv:2408.00118,

  5. [2017]

    LLMs as signal detectors: Sensitivity, bias, and the temperature–criterion analogy

    Jon-Paul Cacioli. LLMs as signal detectors: Sensitivity, bias, and the temperature–criterion analogy. arXiv preprint arXiv:2603.14893,

  6. [2024]

    Table 5: Accuracy by NLP quartile atT=1.0

    A NLP Monotonicity Check Table 5 confirms that NLP is monotonically related to accuracy across all model×dataset conditions atT=1.0, validating its use as a graded evidence variable for Type-2 SDT analysis. Table 5: Accuracy by NLP quartile atT=1.0. All conditions strictly monotonic. Model DatasetQ 1 Q2 Q3 Q4 Llama-3-Instruct TriviaQA 0.179 0.400 0.667 0....

  7. [2026]

    Rescaling confidence: What scale design reveals about LLM metacognition.arXiv preprint arXiv:2603.09309,

    Yuyang Dai. Rescaling confidence: What scale design reveals about LLM metacognition.arXiv preprint arXiv:2603.09309,