REVIEW 3 major objections 4 minor 1 cited by
Like people, large language models are overconfident on hard tasks and underconfident on easy ones.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 13:22 UTC pith:7FYDCXWR
load-bearing objection Solid preregistered multi-model demo of average overconfidence plus a classic hard-easy effect in LLMs, with a useful new continuous-difficulty probe (LifeEval); contamination is real but does not erase the pattern. the 3 major comments →
Confidence Calibration in Large Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Current large language models are, on average, overconfident: stated confidence exceeds accuracy. This overconfidence is moderated by a hard-easy effect—overconfidence grows with task difficulty and flips into underconfidence on easy tasks. LifeEval, a new continuous-difficulty Bayesian-style task grounded in actuarial probabilities, isolates the effect while holding other task features fixed.
What carries the argument
LifeEval: a lifespan-prediction probe that varies sex, minimum age, and accuracy radius so that difficulty is defined as 1 minus the Maximum Achievable Score (the actuarial probability mass inside the optimal radius). Model point estimates and stated confidences are scored against the true conditional probabilities from Social Security period life tables.
Load-bearing premise
The claim that LifeEval cleanly measures difficulty rests on the premise that its actuarial ceiling is free of training-data contamination and other confounds, so that changes in radius truly isolate how models respond to hardness.
What would settle it
Re-run the same models on a LifeEval-style task whose ground-truth distribution is guaranteed never to have appeared in training; if the hard-easy slope of overconfidence versus Maximum Achievable Score disappears or reverses, the isolation claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a preregistered multi-model study of verbal confidence calibration for 11 LLMs (5 reasoning, 6 chat) on six English question sets. Aggregate results show overconfidence (mean stated confidence ~88% vs. accuracy ~79%), moderated by a hard-easy effect: overconfidence is largest on difficult items (LSAT-AR, small-radius LifeEval) while easy items (SciQ, SAT-EN, large-radius LifeEval) produce underconfidence. Reasoning models are better calibrated, less prone to 5% rounding, and more correlated with true probabilities. The novel LifeEval task elicits point estimates of age-at-death plus radius-specific confidence, scored against SSA actuarial probabilities (Eqs. 2-4) to yield a continuous exogenous difficulty measure (1-MAS). ECE (Eq. 9) and overconfidence (Eq. 10) are computed throughout; Appendix G audits SSA contamination and re-runs key plots on the clean subset.
Significance. If robust, the work is a clear contribution: it documents that contemporary LLMs exhibit the classic human hard-easy pattern, with direct implications for user trust, hallucination risk, and possible RLHF-induced overconfidence. Strengths that raise the paper above typical LLM evaluation studies include the preregistration, explicit formulas for ECE and overconfidence, systematic comparison of reasoning vs. chat models, the continuous-difficulty LifeEval construction grounded in public actuarial tables, and the contamination audit that re-analyzes the no-evidence subset. These elements make the central empirical claims falsifiable and relatively careful. The discussion of reflection-style debiasing and the three forms of overconfidence also usefully connects the AI results to the psychological literature.
major comments (3)
- [Appendix G, Tables 4-5, Figure 8] Appendix G / Tables 4-5 / Figure 8: Contamination rates for the strongest reasoning models are high (DeepSeek-R1 71.5%, Gemini-2.5-Pro 71%, GPT-o3 50.1% strong evidence). The no_evidence subset therefore shrinks to N=35-45 for those models. While hard-easy coefficients remain positive, the tiny samples and the selection process itself (clean responses are those that do not cite table values) leave open the possibility that overconfidence on low-MAS items simply reflects inability to compute the actuarial p(k,r|a,s) of Eq. 3 rather than the claimed noisy-confidence mechanism. This undercuts the claim that LifeEval isolates a pure, exogenous difficulty effect (Section 4.5). Either supply CIs/power calculations on the clean subset or substantially qualify the isolation language in the abstract and Discussion.
- [Table 2, Section 7] Table 2 and Section 7: The Hard-Easy regression coefficients (overconfidence on 1-MAS) are presented as point estimates without standard errors, confidence intervals, or any inferential statistics. Reasoning-model averages are modest (0.168) while chat-model averages are large (0.732); without uncertainty quantification it is impossible to judge whether the 'powerful' hard-easy effect is reliable for the better-calibrated models that matter most. Bootstrap or mixed-effects intervals are needed before the central claim can be taken as established.
- [Appendix A, Section 4.5] Appendix A and Section 4.5: LifeEval scoring was changed after preregistration from a binary 'within-radius' indicator to the continuous actuarial probability (Eq. 3). The continuous metric is preferable for calibration analysis, yet it is a material deviation on the paper's flagship task. Sensitivity of the hard-easy slopes and ECE numbers to the original binary scoring should be reported as a robustness check.
minor comments (4)
- [Introduction] Introduction (p. 2): typographical error 'based on of quantitative elements'.
- [Figures 2-4, 8-12] Figures 2-4 and 8-12 lack error bars or confidence bands; given the varying N across models and contamination strata, visual uncertainty would help readers.
- [Table 1] Table 1 and Section 4: HaluEval answer field is listed as N/A; a one-sentence clarification that the model only supplies confidence (not an answer) would improve readability.
- [Section 4.5 / Table 2] Notation for Maximum Achievable Score (MAS) appears first in the LifeEval results without a formal definition until later; move the definition earlier or add a short glossary.
Circularity Check
No circularity: purely empirical comparison of stated confidence against external accuracy or actuarial ground truth, with no fitted parameters or self-referential definitions in the metrics.
full rationale
The paper's central claims (average overconfidence of ~9%, moderated by a hard-easy effect) rest on direct, non-circular comparisons: stated confidence (prompted numerical probabilities, normalized if needed via Eq. 5) versus observed accuracy (Eq. 6) or actuarial success probability (Eqs. 1-4 from public SSA Period Life Tables). Overconfidence is defined as conf(Q) - acc(Q) (Eq. 10); ECE is the standard binned absolute difference (Eq. 9). LifeEval difficulty is 1 - MAS, where MAS is the maximum probability mass inside the optimal radius under the external life tables (Figure 5, Eq. 3); this is independent of any model output and is not estimated from the LLMs. The hard-easy regression simply correlates this exogenous difficulty with the already-computed overconfidence. No parameters are fitted to the evaluation data and then re-used as 'predictions'; no uniqueness theorems or ansatze are imported via self-citation; related-work citations are to external psychology and calibration literature. Contamination screening (Appendix G) is a post-hoc robustness check, not part of the metric definitions. The derivation chain is therefore self-contained against external benchmarks and exhibits zero circular reduction.
Axiom & Free-Parameter Ledger
free parameters (2)
- ECE bin edges
- LifeEval radii set {1,5,10,20}
axioms (3)
- domain assumption Stated numerical confidence in the JSON field is a faithful report of the model's subjective probability of correctness.
- domain assumption SSA Period Life Tables supply the true conditional probabilities p(k,r|a,s) against which LifeEval confidence is scored.
- ad hoc to paper Maximum Achievable Score (MAS) is a pure exogenous difficulty measure that models can detect.
invented entities (1)
-
LifeEval
independent evidence
read the original abstract
We investigate the calibration of large language models' (LLMs') confidence across diverse tasks. The results of our preregistered study show that the current crop of LLMs are, like people, too sure they are right: confidence exceeds accuracy, on average. Importantly, however, this tendency is moderated by a powerful hard-easy effect, wherein overconfidence is greatest on difficult tests; by contrast, easy tests actually show substantial underconfidence. We develop LifeEval, a test for evaluating model calibration across levels of difficulty.
Figures
Forward citations
Cited by 1 Pith paper
-
Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts
A 4B local LLM hits 67.3% on LEED v4.1 credit screening; a deterministic numeric checker lifts EA-p2 from 50% to 100%, but the full neuro-symbolic pipeline trails at 61.6%.
Reference graph
Works this paper leans on
-
[1]
A review of uncertainty quantification in deep learning: Techniques, applications and challenges.Information Fusion, 76:243–297. ArXiv: 2011.06225. Saleh Afroogh, Ali Akbari, Emmie Malone, Mohammadali Kargar, and Hananeh Alambeigi
Pith/arXiv arXiv 2011
-
[2]
Mind the confidence gap: Overconfidence, calibration, and distractor effects in large language models.Preprint, arXiv:2502.11028. Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova
-
[3]
Boolq: Exploring the surprising difficulty of natural yes/no questions.Preprint, arXiv:1905.10044. A. P. Dawid
Pith/arXiv arXiv 1905
-
[4]
Do llms implicitly determine the suitable text difficulty for users?Preprint, arXiv:2402.14453. Google. 2025a. Gemini 2.5 flash. Google. 2025b. Gemini 2.5 pro. Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger
-
[5]
Irene Hou, Hannah Vy Nguyen, Owen Man, and Stephen MacNeil
On calibration of modern neural networks.Preprint, arXiv:1706.04599. Irene Hou, Hannah Vy Nguyen, Owen Man, and Stephen MacNeil
-
[6]
InProceedings of the 56th ACM Technical Symposium on Computer Science Education V
The evolving usage of genai by computing students. InProceedings of the 56th ACM Technical Symposium on Computer Science Education V . 2, SIGCSE TS 2025, page 1481–1482. ACM. Seonjeong Hwang, Hyounghun Kim, and Gary Geunbae Lee
2025
-
[7]
Can llms estimate cognitive complexity of reading comprehension items?Preprint, arXiv:2510.25064. Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bow...
-
[8]
Language models (mostly) know what they know.arXiv, arXiv:2207.05221. ArXiv:2207.05221 [cs]. 10 Daniel Kahneman. 2011.Thinking fast and slow. Farrar, Straus and Giroux, New York. Citation Key: Kahneman2011. Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang
Pith/arXiv arXiv 2011
-
[9]
Why language models hallucinate. Preprint, arXiv:2509.04664. Gideon Keren
-
[10]
Taming overconfidence in llms: Reward calibration in rlhf.arXiv, arXiv:2410.09724. ArXiv:2410.09724 [cs]. Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen
-
[11]
Sarah Lichtenstein and Baruch Fischhoff
Conftuner: Training large language models to express their confidence verbally.Preprint, arXiv:2508.18847. Sarah Lichtenstein and Baruch Fischhoff
-
[12]
Towards robust mathematical reasoning.Preprint, arXiv:2511.01846. Meta. 2024a. Llama 3.1 70b instruct. Meta. 2024b. Llama 3.1 8b instruct. Don A. Moore
-
[13]
When are bayesian model probabilities overconfident?arXiv:2003.04026. ArXiv: 2003.04026. OpenAI
Pith/arXiv arXiv 2003
-
[14]
Understanding model calibration–a gentle introduction and visual exploration of calibration and the expected calibration error (ece).arXiv preprint arXiv:2501.19047. Philip and Hemang
-
[15]
Web page
Period life table, 2022 (used in the 2025 trustees report). Web page. Presented by the Office of the Chief Actuary; accessed via SSA website. Yoo Yeon Sung, Eve Fleisig, Yu Hou, Ishan Upadhyay, and Jordan Lee Boyd-Graber
2022
-
[16]
Grace: A granular benchmark for evaluating model calibration against human calibration.Preprint, arXiv:2502.19684. H. M. Shadman Tabib and Jaber Ahmed Deedar
-
[17]
Toward trustworthy difficulty assessments: Large language models as judges in programming and synthetic tasks.Preprint, arXiv:2511.18597. Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christo- pher D. Manning
-
[18]
Sahil Tripathi, Md Tabrez Nafis, Imran Hussain, and Jiechao Gao
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback.arXiv preprint arXiv:2305.14975. Sahil Tripathi, Md Tabrez Nafis, Imran Hussain, and Jiechao Gao
-
[19]
The confidence paradox: Can llm know when it’s wrong.arXiv, arXiv:2506.23464. ArXiv:2506.23464 [cs]. Thomas S Wallsten, David V Budescu, and Rami Zwick
-
[20]
Crowdsourcing multiple choice science questions.arXiv, arXiv:1707.06209. ArXiv:1707.06209 [cs]. Jiancong Xiao, Bojian Hou, Zhanliang Wang, Ruochen Jin, Qi Long, Weijie J. Su, and Li Shen
-
[21]
Restoring calibration for aligned large language models: A calibration-aware fine-tuning approach.arXiv, arXiv:2505.01997. ArXiv:2505.01997 [cs]. Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi
-
[22]
Chenjun Xu, Bingbing Wen, Bin Han, Robert Wolfe, Lucy Lu Wang, and Bill Howe
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms.Preprint, arXiv:2306.13063. Chenjun Xu, Bingbing Wen, Bin Han, Robert Wolfe, Lucy Lu Wang, and Bill Howe
-
[23]
Do language models mirror human confidence? exploring psychological insights to address overconfidence in llms.Preprint, arXiv:2506.00582. Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan
-
[24]
Agieval: A human-centric benchmark for evaluating foundation models.Preprint, arXiv:2304.06364. Wanjun Zhong, Siyuan Wang, Duyu Tang, Zenan Xu, Daya Guo, Jiahai Wang, Jian Yin, Ming Zhou, and Nan Duan
-
[25]
Ar-lsat: Investigating analytical reasoning of text.Preprint, arXiv:2104.06598. A Deviations from Pre-Registration • DeepSeek log probabilities.DeepSeek did not provide usable token-level log probabilities, so logprob-based analyses for this model were omitted. • LifeEval scoring rule.ForLifeEval, we scored answers using the conditional true (actuarial) p...
-
[26]
Reasoning
For example: Question: <Question> Options: A) <Option A> B) <Option B> C) <Option C> D) <Option D> E) <Option E> Response: { "Reasoning": "<your step-by-step reasoning>", "Answer": "<A, B, C, D, or E>", "A": <float>, "B": <float>, "C": <float>, "D": <float>, "E": <float> } 18 When answering the question about confidence, give a probability that is an hone...
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.