REVIEW 4 major objections 4 minor 9 cited by
The paper argues that an LLM's confidence is a separate, measurable capacity from its knowledge.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 17:25 UTC pith:ZYKTUYB5
load-bearing objection The v3 abstract and the v1 body are different papers: the abstract's new metric never appears in the body, and the body's central finding is explicitly retracted by the abstract. the 4 major comments →
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that confidence quality and knowledge are separable: a model can be a strong discriminator and a weak monitor, or the reverse. In the body's data, the model with the highest Type-1 sensitivity has the lowest metacognitive efficiency, while a lower-sensitivity model sits near optimal, so two confidence metrics rank the four models in opposite order. Efficiency also differs across knowledge domains—different models are weakest in different subjects—and temperature changes the criterion for reporting confidence without changing how much information confidence contains about correctness for two of the instruction-tuned models. The revised abstract states that these strong em
What carries the argument
Type-2 Signal Detection Theory with meta-d'/d' (M-ratio): token-level normalised log-probability is treated as a graded confidence variable that must discriminate the model's own correct from incorrect answers, and meta-d' is the ideal-observer sensitivity that would produce the observed confidence-by-accuracy contingency table. Dividing by Type-1 d' separates monitoring quality from raw knowledge. The revised abstract replaces this ratio with a model-free information measure (meta-I_2r) for the same purpose, on the ground that open-ended QA lacks the two-alternative decision classical meta-d' assumes; the body defines neither the measure nor the resulting numbers.
Load-bearing premise
The load-bearing premise is that the automated correct/incorrect labels are unbiased; the revised abstract admits that a differential length bias in those labels did not survive human adjudication, and if the labels are wrong the body's efficiency comparisons collapse.
What would settle it
Relabel every trial with the 1,830 human adjudications, recompute metacognitive efficiency by model and domain, and check whether the inverse accuracy-efficiency coupling and the reported z-ROC slope range (0.78–1.18) exist; if the coupling survives, the revised abstract's main correction is false, and if it vanishes, the body's headline results are scorer artifacts.
If this is right
- Evaluation reports should list a metacognitive efficiency metric alongside calibration error, because two models with similar accuracy can have opposite usable-confidence profiles.
- For selective prediction systems, model choice flips depending on whether confidence is scored by its ranking quality or by metacognitive efficiency; the paper's efficiency-based choice outperforms the ranking-based choice at equal coverage.
- Domain-level efficiency can hide in aggregate numbers: a model can be metacognitively blind in one subject while efficient in others, so confidence-based deployment should be evaluated per domain.
- Temperature tuning shifts confidence policy, not capacity, for instruction-tuned models; raising temperature cannot fix a model whose confidence carries little information about correctness.
- The revised abstract's admission that the inverse accuracy-efficiency coupling does not survive human relabelling implies the printed conclusions are conditional on the original automated scorer.
Where Pith is reading between the lines
- If the proposed information measure replaces M-ratio, the framework could extend beyond two-alternative QA to open-ended generation, where classical meta-d' is not defined.
- The admitted grader bias is a warning for any LLM-confidence study that relies on automated exact-match scoring: efficiency results should be validated on a human-adjudicated subset before ranking models.
- The revised abstract's near-perfect rank correlation between metacognitive information and the accuracy gain from abstention, if reproducible, would let practitioners predict coverage-accuracy trade-offs without running the system.
- Closing the gap between the revised abstract and the printed body is a prerequisite for testing the strongest claims; the z-ROC slopes and cross-model spread cited in the abstract cannot currently be verified from the published analyses.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript (v1 body, v3 abstract) applies Type-2 Signal Detection Theory to measure metacognitive efficiency in four LLMs across 224,000 factual QA trials. The body defines the meta-d'/d' ratio (M-ratio, Eq. 1) and reports that M-ratio varies across models, is domain-specific, and is dissociated from temperature effects; it also reports an inversion between AUROC2 and M-ratio rankings. The v3 abstract, however, introduces a different metric (normalized metacognitive information, meta-I_2r) and five findings (a 1.98-fold range, rank correlations of -0.80 and +0.00, z-ROC slopes 0.78-1.18, Science & Technology as the weakest domain for every model, near-flat metacognitive information under temperature, and rho = +1.00 with abstention gains) that appear nowhere in the body. The abstract also states that M-ratio 'is not well defined for open-ended QA' and that the inverse accuracy-efficiency coupling reported in v1/v2 does not survive corrected labelling. The submitted text is internally inconsistent, and the central claim is not checkable from the analyses presented.
Significance. A valid model-free metric of usable confidence for open-ended QA would be a significant contribution: it could replace calibration metrics that conflate accuracy and confidence informativeness. The v3 abstract's claims, if supported, would be important. The pre-registration, public code/data, and use of permutation nulls and bootstrap CIs are strengths that partially offset concerns about reproducibility. However, the body contains none of the v3 abstract's analyses: meta-I_2r is not defined, z-ROC slopes are not reported, and the abstract's own admission that M-ratio is not well defined for open-ended QA invalidates the body's central metric. The admitted differential length bias in the automated scorer and the statement that the headline finding does not survive relabelling further undermine the empirical contributions. As submitted, the manuscript cannot be evaluated as a coherent paper.
major comments (4)
- [Abstract vs. body (§2.2, §4)] The v3 abstract and the v1 body describe different studies. The abstract introduces meta-I_2r and reports findings (factor 1.98, rank correlations -0.80/+0.00, z-ROC slopes 0.78-1.18, Science & Technology as weakest domain for every model, near-flat temperature effect, rho = +1.00 with abstention) that appear nowhere in the body. The body defines only M-ratio (Eq. 1) and reports Tables 2, 6, 7; no meta-I_2r is defined or estimated, and no z-ROC analysis is presented. Moreover, the abstract states M-ratio 'is not well defined for open-ended QA, which lacks a two-alternative Type-1 decision', directly contradicting the body's use of M-ratio as its central metric. The referenced 'version note on page 1' is absent. Without the meta-I_2r analysis, the paper's central claim cannot be checked.
- [Abstract vs. §3.1 correctness scoring] The abstract admits that the automated correctness scorer had a 'differential length bias' and that, once corrected against 1,830 human adjudications, 'the inverse accuracy-efficiency coupling reported in v1 and v2 does not survive relabelling'. The body's results (Tables 2-7, Figures 1-3) are all based on the uncorrected labels using exact match plus difflib.SequenceMatcher >= 0.85 (§3.1). No corrected-label analysis is included anywhere in the body. Thus the paper's own headline finding is acknowledged to be an artifact of the grader, and the submitted body still relies on that grader. This is a load-bearing, unresolved problem.
- [§4.3, Table 3] The claim that AUROC2 and M-ratio rankings are 'fully inverted' is substantially forced by construction and overstated. M-ratio = meta-d'/d' normalizes out Type-1 sensitivity by division, while AUROC2 covaries with d', so a negative association between the two rankings is mathematically expected when d' varies across models. In addition, Table 3 itself shows Llama-3-Base and Gemma tied at M-ratio 1.048, so the ranking is not fully inverted between ranks 2 and 3. The interpretation that these metrics answer different questions is acceptable, but presenting the inversion as an empirical discovery overstates what the normalization already implies.
- [§4.2-§4.6 vs. §5-§6] Reported inferential support for the core hypotheses is weak: H1 is partially supported (one of four models has a CI entirely below 1.0, §4.2), H2 is not met (the pre-registered criterion fails, §4.4), H3 is supported for only two of four models (§4.5), and H4 rests on a single significant pairwise comparison (§4.6). The Discussion and Conclusion nevertheless state these as established findings, including domain-specificity and temperature dissociation. The conclusions should be calibrated to the statistics actually reported; the current text overstates the strength of evidence.
minor comments (4)
- [Abstract] The version note on page 1 is referenced but absent; if the version history is important, it should be included or the reference removed.
- [§4.3, Table 3] Use of 'fully inverted' is misleading given the tie between Llama-3-Base and Gemma. Consider describing the ranking difference more precisely.
- [§5.4] The reference 'Goertz et al. (2024)' is mentioned in the text but not listed in the References section. Please add the full citation.
- [§5.1, footnote 1] The footnote reports selective-prediction accuracy at 50% coverage (Gemma 77.4% vs. Mistral 70.7%) but gives no confidence intervals or method details; this is a key practical claim and should be reported with the same inferential rigor as the rest of the paper.
Circularity Check
No circular step can be exhibited; abstract/body mismatch is a checkability failure, not a circular derivation.
full rationale
The paper's core body-level quantities are empirical. Equation (1) defines M-ratio as meta-d'/d'; the AUROC2-vs-M-ratio inversion in Table 3 is a data-dependent ranking, and the paper itself explains that AUROC2 inherits Type-1 sensitivity (Section 4.3) rather than pretending the inversion is a derived consequence. Self-citations to Cacioli (2026) are used for prior data collection and Type-1 d' validation, but they do not smuggle in the paper's new claims, and no uniqueness theorem is imported. The v3 abstract's meta-I_2r, z-ROC slopes, and abstention-gain correlation are not defined or derived in the body, and the abstract's own admissions about scorer bias and the ill-definedness of M-ratio for open-ended QA undermine the v1 results; however absence of an implementation or derivation is a correctness/reproducibility problem, not a circularity that can be exhibited as an equality between the claim and its inputs. Under the hard rule requiring a specific reduction, no circular step is identifiable.
Axiom & Free-Parameter Ledger
free parameters (7)
- Number of confidence bins K =
4 (robustness: 3, 6)
- Bin-edge quantiles =
{12.5, 25, 37.5, 50, 62.5, 75, 87.5}th percentiles of NLP at T=1.0 per model×dataset
- Correctness threshold (difflib ratio) =
0.85
- Log-linear correction =
+0.5 to all cells (Hautus 1995)
- Equal-variance constraint s =
1
- TOST equivalence bound δ =
0.3
- M-ratio instability exclusion =
|M-ratio| > 10 excluded
axioms (5)
- domain assumption NLP is a monotonically valid graded evidence variable for correctness (functional operationalization of confidence)
- domain assumption An equal-variance Gaussian SDT generative model underlies Type-1 and Type-2 responses
- domain assumption Meta-d'/M-ratio is well defined for single-response open-ended QA
- domain assumption Correctness labels from exact match + SequenceMatcher are valid
- standard math Bootstrap resampling at trial level accounts for the data structure
invented entities (1)
-
Normalized metacognitive information (meta-I_2r)
no independent evidence
read the original abstract
Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate two capacities: how much a model knows (Type-1 accuracy) and how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity). We apply Signal Detection Theory to decompose them, treating token-level normalised log-probability as a graded confidence variable and answer correctness as the state to be discriminated. We characterise the Type-2 ROC of this signal, including its unequal-variance structure via z-ROC analysis, and -- because the meta-d' efficiency ratio is not well defined for open-ended QA, which lacks a two-alternative Type-1 decision -- quantify efficiency with a model-free information measure, normalised metacognitive information (meta-I_2r). Across four LLMs and 224,000 factual QA trials we find: (1) metacognitive information varies by a factor of 1.98 across models and is not predicted by accuracy, the rank correlation being -0.80 on TriviaQA and +0.00 on Natural Questions; (2) the confidence signal has model-specific unequal-variance structure (z-ROC slopes 0.78 to 1.18) invisible to calibration metrics, the slope ordering replicating on NQ; (3) efficiency is domain-specific, weakest in Science & Technology for every model; (4) temperature dissociates accuracy from metacognitive information, which stays near-flat for three of four models while accuracy falls; and (5) metacognitive information tracks the accuracy gain from confidence-based abstention exactly (rho = +1.00) while accuracy does not. All estimates carry permutation nulls and bootstrap confidence intervals. This version (v3) corrects a differential length bias in the automated correctness scorer, validated against 1,830 human adjudications; the inverse accuracy-efficiency coupling reported in v1 and v2 does not survive relabelling. See the version note on page 1.
Figures
Forward citations
Cited by 9 Pith papers
-
Do Coding Agents Understand Least-Privilege Authorization?
Coding agents struggle to infer least-privilege file permissions by omitting needed accesses while granting unused or sensitive ones, but Sufficiency-Tightness Decomposition improves sensitive-task success by up to 15...
-
Before You Interpret the Profile: Validity Scaling for LLM Metacognitive Self-Report
Validity indices adapted from clinical assessment classify four frontier LLMs as construct-level invalid on metacognitive probes, with valid models showing positive item-sensitive confidence (r=.18) while invalid ones...
-
Hypothesis generation and updating in large language models
LLMs exhibit Bayesian-like hypothesis updating with strong-sampling bias and an evaluation-generation gap but generalize poorly outside observed data.
-
Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen
Seven 3-9B instruction-tuned LLMs produce verbal confidence that saturates at high values and fails psychometric validity criteria for Type-2 discrimination under minimal elicitation.
-
MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition
MEDLEY-BENCH reveals an evaluation/control dissociation in AI metacognition where scale improves reflective scoring but not proportional belief revision, with a consistent knowing/doing gap across 35 models.
-
K-Way Energy Probes for Metacognition Reduce to Softmax in Discriminative Predictive Coding Networks
K-way energy probes in discriminative PCNs reduce to a monotone function of the log-softmax margin plus an untrained residual and empirically track below softmax on CIFAR-10.
-
Disambiguating electrical detection of magnetization dynamics in magnetic insulators
Spin pumping and ST-FMR compete with opposite voltage signs in Pt/magnetic-insulator devices; thickness excitation profile and damping, not chirality alone, set which wins.
-
Disambiguating electrical detection of magnetization dynamics in magnetic insulators
Spin pumping and ST-FMR contributions to electrical signals in Pt/magnetic-insulator devices can be separated by geometry and field direction, showing that voltage sign is not a unique indicator of magnon chirality.
-
Quantisation Reshapes the Metacognitive Geometry of Language Models
Quantization restructures domain-level M-ratio metacognitive profiles in LLMs while leaving Type-2 AUROC profiles unchanged.
Reference graph
Works this paper leans on
-
[3]
Code and data:https://anonymous.4open.science/r/sdt_calibration
2Pre-registration:https://osf.io/5q7mt/overview?view_only=bd718de95b6c44ff9c14c1ac424227ba. Code and data:https://anonymous.4open.science/r/sdt_calibration. 9 Stephen M. Fleming. HMeta-d: Hierarchical Bayesian estimation of metacognitive efficiency from confidence ratings.Neuroscience of Consciousness, 2017(1):nix007,
2017
-
[1995]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chap- lot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7B.arXiv preprint arXiv:...
-
[1996]
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. Can LLMs ex- press their uncertainty? an empirical evaluation of confidence elicitation in LLMs.arXiv preprint arXiv:2306.13063,
-
[2003]
Gemma 2: Improving open language models at a practical size.arXiv preprint arXiv:2408.00118,
Gemma Team. Gemma 2: Improving open language models at a practical size.arXiv preprint arXiv:2408.00118,
-
[2017]
LLMs as signal detectors: Sensitivity, bias, and the temperature–criterion analogy
Jon-Paul Cacioli. LLMs as signal detectors: Sensitivity, bias, and the temperature–criterion analogy. arXiv preprint arXiv:2603.14893,
-
[2024]
Table 5: Accuracy by NLP quartile atT=1.0
A NLP Monotonicity Check Table 5 confirms that NLP is monotonically related to accuracy across all model×dataset conditions atT=1.0, validating its use as a graded evidence variable for Type-2 SDT analysis. Table 5: Accuracy by NLP quartile atT=1.0. All conditions strictly monotonic. Model DatasetQ 1 Q2 Q3 Q4 Llama-3-Instruct TriviaQA 0.179 0.400 0.667 0....
arXiv 2020
-
[2026]
Yuyang Dai. Rescaling confidence: What scale design reveals about LLM metacognition.arXiv preprint arXiv:2603.09309,
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.