{"id":"f7269334-7b68-4aa9-b229-274bba4280d0","arxiv_id":"2504.12098","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs are overprecise in numerical interval estimation: hit rates fall far below imposed confidence levels, and interval length does not track the requested confidence.","lead":"This paper measures whether large language models can produce numerical intervals that match an imposed confidence level, such as a 90% range. It finds the models are overprecise: their intervals cover the true answer far less often than the stated confidence, and interval width barely changes with the confidence level.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pooled correlation may mask a within-question confidence–length link, leaving the no-correlation claim unverified.","rationale":"The reader's verdict is CONDITIONAL, and I agree that the paper should not be fully accepted without changes. However, the reader's weakest assumption (contamination from sampling all splits) is not the most load-bearing concern. Contamination would, if anything, likely inflate hit rates on memorized answers, so the observed overprecision at high confidence would persist or worsen on unseen data; it does not threaten the central calibration finding. The more fragile claim is the no-correlation result in Section 5.2, which is one of the paper's four stated insights. The pooled Pearson correlation is computed over clustered data (each question at five confidence levels), and between-question variance in interval length can hide a within-question positive relationship between confidence and interval width. The paper's own Figure 5 demonstrates large between-question and between-dataset length variation, making this concern concrete. Reanalyzing the correlation within questions is a straightforward, decisive check. If the within-question correlation is positive, the paper's conclusion that LLMs cannot adjust interval width to instructed confidence would be wrong, and the overprecision narrative would need to be reframed around hit-rate miscalibration alone. Since the hit-rate evidence is strong, the verdict should remain CONDITIONAL pending this reanalysis, rather than REJECT or UNVERDICTED.","tokens_in":19169,"tokens_out":8371,"duration_ms":89952,"concrete_test":"For each question in each Table 4 model–dataset–prompt cell, compute the Pearson (or Spearman) correlation between the five imposed confidence levels and the corresponding generated interval lengths. Average these per-question correlations and test whether the mean is significantly greater than zero using a one-sample t-test or Wilcoxon signed-rank test. Also fit a linear mixed model with a fixed slope for confidence and a random intercept per question. If the within-question mean correlation or the mixed-model slope is positive and statistically significant for any cell, the reported pooled correlation of ~0 is an aggregation artifact and the Section 5.2 no-correlation claim must be revised or qualified.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 3.3 defines Pearson's correlation between imposed confidence levels and interval lengths, and Table 4 reports near-zero correlations as evidence for 'no correlation between the length of the interval and the imposed confidence level' (Section 5.2). The data, however, have a repeated-measures structure: each question is evaluated at five confidence levels, and interval length varies strongly by question and by answer scale (the paper's own ILS analysis in Appendix D shows large between-question and between-dataset differences). The pooled correlation across all question–confidence pairs conflates between-question variation with the within-question effect that the authors intend to measure. If a genuine within-question effect exists—wider intervals for higher confidence—it can be swamped by the between-question variance in interval length, yielding a pooled correlation near zero. The paper therefore does not establish that interval length is unrelated to imposed confidence; it only shows that the pooled correlation is weak. The correct analysis would estimate the within-question slope via a mixed-effects model with a random intercept per question, or by computing per-question correlations and testing their mean. Without such an analysis, the central claim in Section 5.2 remains unsupported, even though the hit-rate calibration failure at high confidence (Tables 4) is credible.","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-16T12:38:13.866517+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}