{"id":"03d0e69c-79a7-4c69-9454-de6a6625d59d","arxiv_id":"2603.25112","paper_version":3,"verdict":"REJECT","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":7,"one_line_summary":"The claimed finding — that a new model-free 'metacognitive information' metric varies across LLMs by 1.98×, is uncorrelated with accuracy, and exactly tracks abstention gains — appears only in the abstract, which contradicts the body's M-ratio analysis and retracts that analysis's headline results.","lead":"The paper proposes signal-detection measures of how well an LLM's confidence tracks whether its answer is correct, and claims such 'metacognitive efficiency' varies across models in ways standard calibration metrics miss. The abstract and the body, however, describe different metrics and contradict each other on the main findings, and the abstract's headline results are not present in the body text.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The v3 abstract and v1 body are different papers: meta-I_2r, z-ROC slopes, and the abstract's five findings appear nowhere in the body, which instead defines only M-ratio and contradicts those findings. The central claim is not checkable.","rationale":"The reader's verdict of REJECT is, if anything, understated: the manuscript does not contain one coherent central claim. The body is a v1 preprint about M-ratio; the abstract is a v3 statement about meta-I_2r with different numbers and even a different domain ranking. The reader's weakest assumption focused on label bias, which is important, but the more load-bearing problem is that the v3 claims are not merely unverified — they are absent from the body. The abstract's own admission that M-ratio is not well defined for open-ended QA invalidates the body's headline metric, and the promised definition of the replacement metric is missing. This is not a matter of outside-consensus disagreement; it is an internal inconsistency that prevents any checkable scientific assertion. The v3 abstract also acknowledges the scorer bias that the reader flagged, and the body still uses the biased scorer. While the body's v1 findings (e.g., Mistral's low M-ratio) might be internally consistent under v1's assumptions, that v1 story is explicitly retracted by the abstract's relabelling statement. A rejection is appropriate because no version of the submitted manuscript supports the claims on its own abstract. I agree only partially with the reader's weakest_assumption because the decisive failure is the version mismatch and missing metric, not just the label bias — though both are real and mutually reinforcing.","tokens_in":10978,"tokens_out":4806,"duration_ms":55923,"concrete_test":"Obtain the promised v3 version note/full text and verify that it actually defines meta-I_2r, reports z-ROC slopes, and reproduces the v3 abstract's five quantitative claims (1.98-fold range, rho = -0.80/+0.00, slopes 0.78–1.18, Science & Technology weakest for all models, rho = +1.00 with abstention) from the public code and human-adjudicated labels. If the note is absent, or if any of these quantities does not appear or is not reproduced, the abstract's central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The manuscript's own indexing abstract (v3) and body (v1) do not describe the same study. The body defines only M-ratio (Eq. 1) and reports M-ratios between 0.852 and 1.048, with model-specific weakest domains (Table 6: Mistral is weakest in Arts & Literature, 0.677; Llama-3-Base in History & Politics, 0.894) and temperature effects that vary by model (Table 7: Llama-3-Base d' rises from 1.174 to 1.407, M rises from 0.946 to 1.048). The v3 abstract instead claims a different metric, meta-I_2r, and five findings — a 1.98-fold range in metacognitive information, rank correlations of -0.80 and +0.00 with accuracy, z-ROC slopes 0.78–1.18, Science & Technology as the weakest domain for every model, near-flat metacognitive information under temperature for three of four models, and rho = +1.00 with abstention gains. None of these quantities or results is derived or reported in the body. The abstract itself states that M-ratio 'is not well defined for open-ended QA', so the body's central metric is explicitly declared invalid, while the replacement metric is never defined. The referenced 'version note on page 1' is absent. Furthermore, the abstract admits the automated scorer used in the body (§3.1) had a differential length bias, and that re-labelling with 1,830 human adjudications removes the v1/v2 inverse coupling; the body contains no corrected-label analysis. The central claim therefore cannot be checked against any reproducible analysis in this manuscript.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript (v1 body, v3 abstract) applies Type-2 Signal Detection Theory to measure metacognitive efficiency in four LLMs across 224,000 factual QA trials. The body defines the meta-d'/d' ratio (M-ratio, Eq. 1) and reports that M-ratio varies across models, is domain-specific, and is dissociated from temperature effects; it also reports an inversion between AUROC2 and M-ratio rankings. The v3 abstract, however, introduces a different metric (normalized metacognitive information, meta-I_2r) and five findings (a 1.98-fold range, rank correlations of -0.80 and +0.00, z-ROC slopes 0.78-1.18, Science & Technology as the weakest domain for every model, near-flat metacognitive information under temperature, and rho = +1.00 with abstention gains) that appear nowhere in the body. The abstract also states that M-ratio 'is not well defined for open-ended QA' and that the inverse accuracy-efficiency coupling reported in v1/v2 does not survive corrected labelling. The submitted text is internally inconsistent, and the central claim is not checkable from the analyses presented.","tokens_in":11213,"tokens_out":4505,"duration_ms":47048,"significance":"A valid model-free metric of usable confidence for open-ended QA would be a significant contribution: it could replace calibration metrics that conflate accuracy and confidence informativeness. The v3 abstract's claims, if supported, would be important. The pre-registration, public code/data, and use of permutation nulls and bootstrap CIs are strengths that partially offset concerns about reproducibility. However, the body contains none of the v3 abstract's analyses: meta-I_2r is not defined, z-ROC slopes are not reported, and the abstract's own admission that M-ratio is not well defined for open-ended QA invalidates the body's central metric. The admitted differential length bias in the automated scorer and the statement that the headline finding does not survive relabelling further undermine the empirical contributions. As submitted, the manuscript cannot be evaluated as a coherent paper.","major_comments":[{"comment":"The v3 abstract and the v1 body describe different studies. The abstract introduces meta-I_2r and reports findings (factor 1.98, rank correlations -0.80/+0.00, z-ROC slopes 0.78-1.18, Science & Technology as weakest domain for every model, near-flat temperature effect, rho = +1.00 with abstention) that appear nowhere in the body. The body defines only M-ratio (Eq. 1) and reports Tables 2, 6, 7; no meta-I_2r is defined or estimated, and no z-ROC analysis is presented. Moreover, the abstract states M-ratio 'is not well defined for open-ended QA, which lacks a two-alternative Type-1 decision', directly contradicting the body's use of M-ratio as its central metric. The referenced 'version note on page 1' is absent. Without the meta-I_2r analysis, the paper's central claim cannot be checked.","section":"Abstract vs. body (§2.2, §4)"},{"comment":"The abstract admits that the automated correctness scorer had a 'differential length bias' and that, once corrected against 1,830 human adjudications, 'the inverse accuracy-efficiency coupling reported in v1 and v2 does not survive relabelling'. The body's results (Tables 2-7, Figures 1-3) are all based on the uncorrected labels using exact match plus difflib.SequenceMatcher >= 0.85 (§3.1). No corrected-label analysis is included anywhere in the body. Thus the paper's own headline finding is acknowledged to be an artifact of the grader, and the submitted body still relies on that grader. This is a load-bearing, unresolved problem.","section":"Abstract vs. §3.1 correctness scoring"},{"comment":"The claim that AUROC2 and M-ratio rankings are 'fully inverted' is substantially forced by construction and overstated. M-ratio = meta-d'/d' normalizes out Type-1 sensitivity by division, while AUROC2 covaries with d', so a negative association between the two rankings is mathematically expected when d' varies across models. In addition, Table 3 itself shows Llama-3-Base and Gemma tied at M-ratio 1.048, so the ranking is not fully inverted between ranks 2 and 3. The interpretation that these metrics answer different questions is acceptable, but presenting the inversion as an empirical discovery overstates what the normalization already implies.","section":"§4.3, Table 3"},{"comment":"Reported inferential support for the core hypotheses is weak: H1 is partially supported (one of four models has a CI entirely below 1.0, §4.2), H2 is not met (the pre-registered criterion fails, §4.4), H3 is supported for only two of four models (§4.5), and H4 rests on a single significant pairwise comparison (§4.6). The Discussion and Conclusion nevertheless state these as established findings, including domain-specificity and temperature dissociation. The conclusions should be calibrated to the statistics actually reported; the current text overstates the strength of evidence.","section":"§4.2-§4.6 vs. §5-§6"}],"minor_comments":[{"comment":"The version note on page 1 is referenced but absent; if the version history is important, it should be included or the reference removed.","section":"Abstract"},{"comment":"Use of 'fully inverted' is misleading given the tie between Llama-3-Base and Gemma. Consider describing the ranking difference more precisely.","section":"§4.3, Table 3"},{"comment":"The reference 'Goertz et al. (2024)' is mentioned in the text but not listed in the References section. Please add the full citation.","section":"§5.4"},{"comment":"The footnote reports selective-prediction accuracy at 50% coverage (Gemma 77.4% vs. Mistral 70.7%) but gives no confidence intervals or method details; this is a key practical claim and should be reported with the same inferential rigor as the rest of the paper.","section":"§5.1, footnote 1"}],"recommendation":"reject","confidential_remarks":"The abstract/body mismatch suggests a version-control failure. Even taken on its own terms, the body's results do not support the claims in the abstract, and the admitted scorer bias invalidates the original headline. A major rewrite with new analyses on corrected labels would be needed before this could be considered; as submitted, the manuscript is not coherent enough to publish."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the manuscript as posted is two different papers. The arXiv abstract (v3) announces a new metric, meta-I_2r, and five findings that do not appear anywhere in the body. The body (v1) is a pre-registered Type-2 SDT analysis using M-ratio. The abstract explicitly retracts the body's headline result and says M-ratio \"is not well defined for open-ended QA.\" So the central claim is not checkable from this text.\n\nWhat is genuinely useful: applying meta-d'/M-ratio to internal token log-probabilities is a legitimate extension of an established framework, and the body executes it carefully in places. There is real work here: 224K trials, four models, pre-registered hypotheses, bootstrap CIs, binning robustness checks, and a monotonicity validation of NLP as a graded confidence signal. The dissociation between d' and meta-d' for Mistral and Gemma under temperature is the kind of finding that would interest people working on selective prediction. The AUROC2 vs M-ratio inversion is less novel than it looks because M-ratio normalizes out Type-1 sensitivity by construction, but the empirical demonstration is still useful.\n\nThe soft spots are not minor. The version mismatch alone is disqualifying for this submission. The v3 abstract reports rank correlations across four models (n=4), which carry no statistical weight. The abstract admits the automated scorer had a differential length bias, corrected with 1,830 human adjudications, and that the inverse accuracy-efficiency coupling disappears — meaning the body's main result was an artifact of the grader. The body's own H2 (domain-specificity) failed by the pre-registered criterion. M-ratio > 1 for two models is theoretically suspect given the sampling-bottleneck discussion. The abstract's claims contradict the body's tables: the weakest domain is not Science & Technology for every model, and temperature effects vary by model.\n\nFairly: the author knows the score. The v3 abstract is an honest retraction, and the limitations section is thoughtful. But a reader cannot verify the new claims because meta-I_2r is never defined or derived anywhere in the manuscript.\n\nMy recommendation: don't referee this version. Ask the author to fix the version mismatch, define and derive meta-I_2r, and re-present the corrected analyses. The underlying question — whether there is a model-free, accuracy-independent measure of usable confidence — is worth pursuing, and this author has the data and the willingness to correct course. But the manuscript as it stands is not a reliable basis for evaluation.","headline":"The v3 abstract and the v1 body are different papers: the abstract's new metric never appears in the body, and the body's central finding is explicitly retracted by the abstract.","tokens_in":11922,"tokens_out":2903,"would_cite":false,"duration_ms":30808,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that an LLM's confidence is a separate, measurable capacity from its knowledge.","keywords":["LLM confidence","metacognitive efficiency","meta-d' / M-ratio","Type-2 signal detection theory","calibration","selective prediction","token log-probability","factual question answering"],"falsifier":"Relabel every trial with the 1,830 human adjudications, recompute metacognitive efficiency by model and domain, and check whether the inverse accuracy-efficiency coupling and the reported z-ROC slope range (0.78–1.18) exist; if the coupling survives, the revised abstract's main correction is false, and if it vanishes, the body's headline results are scorer artifacts.","tokens_in":10639,"feed_emoji":"🧠","tokens_out":9149,"duration_ms":91177,"temperature":0.7,"pith_summary":"Current confidence metrics conflate two capacities: whether a model produces correct answers and whether its confidence signal marks the answers it gets wrong. The paper applies Type-2 signal detection theory to token-level log-probabilities, measuring metacognitive efficiency as the ratio of meta-d' to d' across four open-weight models and 224,000 factual QA trials. It claims efficiency is independent of accuracy—the model with the best discrimination has the least usable confidence—varies by knowledge domain, and is dissociated from temperature, which shifts confidence placement while leaving the information content near-flat for three of four models. A revised abstract makes a further claim the printed body does not support: that after correcting a grading bias against human adjudications, the inverse accuracy-efficiency coupling disappears and efficiency should be measured by a model-free information measure. If the core claim holds, model selection for tasks that rely on confidence should use a monitoring metric rather than calibration error.","feed_headline":"How well an LLM knows it's right differs from accuracy","feed_subtitle":"Type-2 signal-detection analysis of 224,000 QA trials gives a usable-confidence metric calibration scores miss.","key_machinery":"Type-2 Signal Detection Theory with meta-d'/d' (M-ratio): token-level normalised log-probability is treated as a graded confidence variable that must discriminate the model's own correct from incorrect answers, and meta-d' is the ideal-observer sensitivity that would produce the observed confidence-by-accuracy contingency table. Dividing by Type-1 d' separates monitoring quality from raw knowledge. The revised abstract replaces this ratio with a model-free information measure (meta-I_2r) for the same purpose, on the ground that open-ended QA lacks the two-alternative decision classical meta-d' assumes; the body defines neither the measure nor the resulting numbers.","core_discovery":"The central claim is that confidence quality and knowledge are separable: a model can be a strong discriminator and a weak monitor, or the reverse. In the body's data, the model with the highest Type-1 sensitivity has the lowest metacognitive efficiency, while a lower-sensitivity model sits near optimal, so two confidence metrics rank the four models in opposite order. Efficiency also differs across knowledge domains—different models are weakest in different subjects—and temperature changes the criterion for reporting confidence without changing how much information confidence contains about correctness for two of the instruction-tuned models. The revised abstract states that these strong em","pith_inferences":["If the proposed information measure replaces M-ratio, the framework could extend beyond two-alternative QA to open-ended generation, where classical meta-d' is not defined.","The admitted grader bias is a warning for any LLM-confidence study that relies on automated exact-match scoring: efficiency results should be validated on a human-adjudicated subset before ranking models.","The revised abstract's near-perfect rank correlation between metacognitive information and the accuracy gain from abstention, if reproducible, would let practitioners predict coverage-accuracy trade-offs without running the system.","Closing the gap between the revised abstract and the printed body is a prerequisite for testing the strongest claims; the z-ROC slopes and cross-model spread cited in the abstract cannot currently be verified from the published analyses."],"forward_implications":["Evaluation reports should list a metacognitive efficiency metric alongside calibration error, because two models with similar accuracy can have opposite usable-confidence profiles.","For selective prediction systems, model choice flips depending on whether confidence is scored by its ranking quality or by metacognitive efficiency; the paper's efficiency-based choice outperforms the ranking-based choice at equal coverage.","Domain-level efficiency can hide in aggregate numbers: a model can be metacognitively blind in one subject while efficient in others, so confidence-based deployment should be evaluated per domain.","Temperature tuning shifts confidence policy, not capacity, for instruction-tuned models; raising temperature cannot fix a model whose confidence carries little information about correctness.","The revised abstract's admission that the inverse accuracy-efficiency coupling does not survive human relabelling implies the printed conclusions are conditional on the original automated scorer."],"fun_headline_variants":["LLM accuracy and confidence quality are uncoupled","Higher accuracy doesn't mean better self-knowledge","LLM metacognition lags behind its knowledge","Confidence quality in LLMs: a separate dimension","Why LLMs can be right but not know they're right"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the automated correct/incorrect labels are unbiased; the revised abstract admits that a differential length bias in those labels did not survive human adjudication, and if the labels are wrong the body's efficiency comparisons collapse.","fun_headline_variants_meta":{"raw":{"variants":["LLM accuracy and confidence quality are uncoupled","Higher accuracy doesn't mean better self-knowledge","LLM metacognition lags behind its knowledge","Confidence quality in LLMs: a separate dimension","Why LLMs can be right but not know they're right"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000296,"raw_usage":{"total_tokens":1625,"prompt_tokens":888,"completion_tokens":737,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":662}},"tokens_in":632,"tokens_out":737,"duration_ms":8489,"temperature":1.0,"reasoning_tokens":662,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T17:25:50.869910+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Relabel every trial with the 1,830 human adjudications, recompute metacognitive efficiency by model and domain, and check whether the inverse accuracy-efficiency coupling and the reported z-ROC slope range (0.78–1.18) exist; if the coupling survives, the revised abstract's main correction is false, and if it vanishes, the body's headline results are scorer artifacts.","supporting_citations":[],"review_version":2}