{"id":"af39bb57-2f17-43fe-852c-a1e73a9ca3bc","arxiv_id":"2506.07461","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM uncertainty quantification should be judged by whether it improves real human decisions, not by calibration scores on trivia benchmarks.","lead":"This position paper argues that current ways of measuring uncertainty in large language models focus on benchmarks and calibration scores that do not tell us whether users actually make better decisions. It reviews 40 LLM uncertainty methods and finds most are tested on trivia-style tasks, ignoring the messy, high-stakes, or genuinely random situations where people need help.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's quantitative evidence is not currently auditable: it claims 40 methods but Appendix B annotates only 37, and Table 1's prose conflicts with its own checkmarks on NQ-Open's C3 flag and on the C3 total; the prevalence claims depend on this corpus.","rationale":"The paper's central assertion is that the field's prevailing UQ evaluation practices are insufficient for human users. To support that, it offers a structured census: 40 methods, 22 benchmarks, and Table 1 annotations. This census is the paper's main novel evidence; the rest of the argument synthesizes prior human-AI collaboration results. The 40-vs-37 mismatch means the census is not fully specified, and the Table 1 inconsistencies mean the annotation protocol is not yet stable enough to support exact counts. These are not stylistic issues: \"only 2 benchmarks satisfy C1–C3\" and \"10 out of 19 supervised methods ignore distribution shift\" are the concrete facts from which the paper generalizes. If the missing three papers include human-uplift studies or distribution-shift evaluations, or if the annotation rubric yields different counts under independent application, the prevalence claims weaken. The normative recommendations could still stand as a call for more human studies, but the paper's evidence would need to be qualified. This matches the reader's conditional verdict: the argument is reasonable and worth publishing as a call to action, but the quantitative backbone should be corrected and made auditable before the evidence is treated as definitive. I therefore keep the verdict conditional and agree with the reader's identification of the annotation corpus as a key weakness, while sharpening it to specific internal inconsistencies rather than only subjectivity.","tokens_in":24779,"tokens_out":8866,"duration_ms":100851,"concrete_test":"Reconstruct the survey corpus from Shorinwa et al. (2024) and Huang et al. (2024) using the paper's stated scope criteria; identify the three missing papers and re-run the count-based claims (supervised methods evaluated under distribution shift, methods using ECE, any human uplift studies). Independently recompute the C1–C4 annotations of Table 1 from the benchmark definitions with two annotators and report inter-annotator agreement, then reconcile the C3 total and the NQ-Open row. If the corrected counts change the headline percentages (e.g., fewer than \"only 2\" C1–C3 benchmarks, or more than 10/19 supervised methods testing distribution shift), the quantitative support for the central claim is weaker than presented.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical generalization about what \"prevailing practices\" in LLM UQ evaluation look like, so the survey corpus and benchmark annotations are load-bearing. Two concrete defects undermine that evidence. First, the paper states it analyzes 40 LLM UQ method papers (Abstract, Section 2), but Appendix B contains only 37 annotated entries (P1–P37); the three missing papers are never identified, so statistics such as \"10 out of 19 supervised parameter-learning methods\" tested under distribution shift (Section 4.2) and \"16 papers used ECE\" (Appendix A) cannot be reproduced or checked for selection bias. Second, the ecological-validity table is internally inconsistent: Section 3.2 says two benchmarks (Natural Questions and NQ-Open) satisfy C1–C3, but Table 1 marks NQ-Open's C3 as ✗; the text also says 9 of 22 benchmarks satisfy C3, while counting the ✓ entries in Table 1 gives 10. These discrepancies make the precise percentages (27.3%, \"only 2\") unreliable. The qualitative argument may survive, but the paper's distinctive empirical contribution—the structured demonstration that current benchmarks are not human-centered—rests on a corpus and annotation set that need to be corrected and independently verified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that the NLP community's prevailing practices for evaluating LLM uncertainty quantification (UQ) are not sufficient for building UQ that benefits human users in real-world decision-making. The authors support this with a survey of 40 LLM UQ method papers and 22 benchmarks, and they identify three barriers: benchmarks with low ecological validity, a narrow focus on epistemic uncertainty, and metrics not tied to downstream utility. For each barrier, they propose concrete recommendations and research directions, and they discuss two alternative views in Section 6. The paper's main empirical contribution is the structured annotation of benchmarks against four criteria (C1–C4) adapted from human-centered XAI work, together with frequency counts such as 'only 2 benchmarks satisfy C1–C3' and '10 out of 19 supervised parameter-learning methods tested under distribution shift.'","tokens_in":25017,"tokens_out":6363,"duration_ms":67447,"significance":"If the empirical survey is corrected and made auditable, the paper could serve as a useful programmatic statement for the LLM UQ community. Its strengths are that it makes falsifiable quantitative claims about the field's evaluation practices, imports concrete criteria from the XAI literature rather than relying on vague objections, and pairs each critique with actionable recommendations. The paper does not contain fitted parameters or circular derivations; it leans on external published results, including two from the authors' own group, which is appropriate for a position paper. The central qualitative message—that calibration improvements on QA benchmarks do not automatically translate into better human-LLM collaboration—is well supported by the cited classical UQ and HCI literature.","major_comments":[{"comment":"The paper states that it analyzes 40 LLM UQ method papers and that Appendix B provides an annotation of each selected paper, but Appendix B contains only 37 entries (P1–P37). This discrepancy is load-bearing because Section 4.2's '10 out of 19 supervised, parameter-learning methods' statistic and Appendix A's '16 papers used ECE' count are computed on the survey corpus. The authors should either add the missing three annotations or revise the stated corpus size to 37 and recompute all dependent statistics; otherwise the prevalence claims cannot be reproduced or checked for selection bias.","section":"§2 and Appendix B"},{"comment":"The text says that exactly two benchmarks satisfy C1–C3 and names Natural Questions and NQ-Open, but Table 1 marks NQ-Open's C3 as ✗. Under the stated criterion, only Natural Questions satisfies C1–C3. I verified that the C3 count in Table 1 is 9, matching the text, so that specific discrepancy is not present; however, the NQ-Open flag contradicts the 'only 2 benchmarks' sentence, which is the headline ecological-validity result. The authors should correct the table or the text and ensure that the 27.3% C1 percentage and the 'only 2' claim are computed from the final annotation.","section":"§3.2 and Table 1"},{"comment":"The C1–C4 annotations are the sole quantitative basis for the ecological-validity claims, but the paper provides no annotation protocol, no inter-annotator agreement, and no released annotation artifact. Several judgments are borderline and contestable (e.g., WebQA receives C1 but not C2; NQ-Open is treated as satisfying C3 in the text but not in the table). Without a rubric or a second annotator, the 27.3% and 'only 2' statements are not independently checkable. The authors should provide the annotation guidelines, ideally with agreement statistics or a public artifact.","section":"§3.2"},{"comment":"Restricting the benchmark pool to benchmarks used by at least two of the surveyed papers may bias the analysis toward established QA and commonsense benchmarks and away from newer, more ecological tasks. For example, FolkTexts is cited in Section 4.1 as a recommended aleatoric-uncertainty benchmark but does not appear in Table 1. Because the paper's claim that 'the majority of LLM UQ methods are evaluated on only factual QA or commonsense reasoning' is computed on this restricted set, the authors should justify the two-paper rule with a sensitivity analysis or report the full pool of benchmarks.","section":"Footnote 2 and §3"}],"minor_comments":[{"comment":"There is a missing space in 'community’sprevailing practices'.","section":"§1.1"},{"comment":"The phrase 'estimate uncertain when the set of possible decisions is not enumerated' should be 'estimate uncertainty when'.","section":"§3.3"},{"comment":"The sentence 'the clinician is still required to making a treatment decision' contains a subject-verb error; it should be 'required to make a treatment decision'.","section":"§4.1"},{"comment":"The author name 'Vodrahalli' is typeset as 'V odrahalli' in the text and references, and the reference format is inconsistent with the rest of the bibliography.","section":"§5.1 and References"},{"comment":"The citation style for the same research group is inconsistent: 'Corvelo Benz and Rodriguez (2023)' in the text but 'Corvelo Benz and Gomez Rodriguez (2025)' elsewhere, with the reference list using both forms.","section":"§5.1 and References"},{"comment":"Several dataset names contain spacing artifacts ('SW AG', 'HellaSW AG') that should be 'SWAG' and 'HellaSwag' for consistency with the cited papers.","section":"Table 1 and Appendix B"},{"comment":"In recommendation R7, 'Brier scorecorrelate' is missing a space; it should read 'Brier score correlate'.","section":"§5.2"},{"comment":"The sentence 'We argue that the LLM UQ methods primarily evaluate on benchmarks with low ecological validity' should read 'the LLM UQ methods are primarily evaluated'.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's qualitative argument is defensible and likely to be of interest to the cs.CL community, but the quantitative survey evidence needs careful correction and auditability before publication. The 40-versus-37 discrepancy and the Table 1 contradiction are simple to fix in principle, but they currently undermine the paper's distinctive empirical contribution. I encourage the editor to request the annotation artifact and the missing paper list as part of the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: this is a solid position paper whose new empirical wrapper is under-audited. The qualitative case—that LLM UQ evaluation is too far from human utility—is not brand new, and the paper does a good job consolidating prior critiques (De Vries et al., Błasiok and Nakkiran, Corvelo Benz et al.) into three sharp barriers. What is new is the structured annotation of 22 benchmarks and the concrete counts: 6/22 pass C1, 15/22 are multiple choice, only 2 pass C1–C3. That artifact alone makes the paper worth engaging.\n\nThe recommendations (R1–R8) are actionable and mostly sensible. The paper also resists the obvious strawman with a reasonable alternative-views section.\n\nSoft spots, in order of softness. First, the survey corpus is not auditable as written. The abstract and Section 2 say 40 methods, but Appendix B annotates only 37 (P1–P37). The missing three are never identified. Since the headline counts—10/19 supervised methods tested under OOD, 16 papers using ECE—depend on that corpus, the authors need to ship the full list and reconcile the count. Second, Table 1 has an internal inconsistency: Section 3.2 says Natural Questions and NQ-Open both satisfy C1–C3, but Table 1 marks NQ-Open’s C3 as ✗. The \"9 of 22 satisfy C3\" count does match the table as printed, so this is localized, but it is exactly the kind of error that makes a reviewer doubt the other annotations. Third, the C1–C4 annotations are one-off subjective judgments with no reliability data. That is not fatal for a position paper, but it means the percentages should be read as illustrative, not definitive. The two-paper inclusion rule also likely biases the benchmark pool toward QA/commonsense, which the paper could acknowledge more directly.\n\nThe central argument survives these problems. The paper is honest about its scope, and the self-citations (Hansen et al., Srinivasan and Thomason) are to external published findings, not circular moves.\n\nWho this is for: anyone working on LLM UQ evaluation, and people who care about how the field defines progress. It is a call to action, not a technical result. I think it deserves a serious referee: send it out, but the authors should be asked to fix the corpus count, the NQ-Open flag, and ideally provide the full annotation with criteria examples before the evidence is treated as definitive.\n\nMy recommendation: engage with it, conditional on those corrections.","headline":"A well-argued position paper with a useful but under-audited benchmark annotation; the qualitative case holds, but the 40/37 count and Table 1 inconsistencies need fixing before the evidence is treated as definitive.","tokens_in":25603,"tokens_out":3002,"would_cite":true,"duration_ms":31810,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM uncertainty scores don't make LLMs safer for users, a 40-method survey argues.","keywords":["LLM uncertainty quantification","calibration","ecological validity","human-LLM collaboration","aleatoric uncertainty","distribution shift","benchmark evaluation","human uplift"],"falsifier":"A human-subject study in which presenting calibrated LLM confidence scores to users produces a measurable, consistent improvement in decision accuracy on a realistic task, while uncalibrated scores do not, would directly contradict the claim that calibration metrics are unconnected to downstream utility. Alternatively, an independent re-annotation of the 22 benchmarks that placed the majority in different C1 or C2 categories would weaken the ecological-validity counts.","tokens_in":24567,"feed_emoji":"🤝","tokens_out":4814,"duration_ms":51099,"temperature":0.7,"pith_summary":"This position paper argues that the way uncertainty quantification (UQ) for large language models (LLMs) is currently evaluated does not serve the stated goal of helping human users decide when to trust model outputs. Reviewing 40 LLM UQ methods and the 22 benchmarks those methods rely on, it identifies three systemic problems: the benchmarks are not representative of real-world tasks, they ignore aleatoric uncertainty and distribution shift, and the metrics being optimized (like ECE and Brier score) have not been shown to improve human decision making. The paper's position is that the field should shift from hill-climbing calibration scores to a human-centered approach that tests whether uncertainty information actually gives users measurable benefit. This matters because LLMs are already deployed as decision aids in high-stakes settings, where a well-calibrated confidence score that does not change user behavior offers little protection.","feed_headline":"Uncertainty scores don't make LLMs safer for users","feed_subtitle":"A 40-method survey says benchmarks and metrics miss what real decision-makers need.","key_machinery":"The argument is carried by a systematic annotation exercise: the authors take 40 LLM UQ papers from two recent surveys, extract the 22 benchmarks used by at least two papers, and score each benchmark against four criteria (C1–C4) adapted from explainable-AI research: connection to a real task, realistic inputs, genuine difficulty for people, and potential for harm from bad decisions. They also classify each benchmark by uncertainty type (epistemic, aleatoric, distributional) and each method by supervision and evaluation metric. This yields the quantitative counts, such as 15 of 22 benchmarks being multiple choice and only 2 of 22 satisfying C1–C3, that anchor the three claimed barriers.","core_discovery":"On the paper's own terms, the central claim is that the NLP community's prevailing evaluation practices for LLM UQ methods are insufficient to benefit human users in real-world settings. The paper supports this with a survey of 40 UQ method papers, finding that most benchmarks are factual QA or commonsense tasks, that only two benchmarks satisfy all the adopted ecological-validity criteria (Natural Questions and NQ-Open), that only one benchmark intentionally contains aleatoric uncertainty (AmbigQA), that fewer than half of supervised methods are tested under distribution shift, and that almost all papers report calibration metrics without any human-uplift study. It concludes that better ECE scores on these benchmarks do not automatically translate into better human-LLM collaboration.","pith_inferences":["If the paper is right, the same critique likely applies to other AI-assistance fields that optimize calibration metrics without measuring user outcomes, so the C1–C4 screen could be reused as a pre-filter for any human-AI collaboration benchmark.","A testable extension would be to correlate ECE improvements with human uplift across the few existing studies that include human evaluations, to see whether any threshold of calibration quality actually predicts better joint decisions.","The paper's own recommendation implies a concrete measurable goal: an LLM UQ method should be judged by the increase in joint human-AI decision accuracy per unit of user trust expenditure, not by calibration error alone."],"forward_implications":["LLM UQ papers should report results on tasks with intentional aleatoric uncertainty, such as datasets with multiple gold labels or conflicting outcomes, rather than only single-answer QA datasets.","Supervised UQ methods should be tested under distribution shift as a standard requirement, since calibrated parameters learned on one distribution may not transfer to new query distributions.","Calibration metrics like ECE and Brier score should be supplemented or replaced by metrics validated against human decision performance, and smECE should be preferred over ECE for evaluation sets smaller than about 5,000 examples.","The community should invest in human-uplift studies that compare decisions made with and without uncertainty information, including non-numeric presentation schemes such as hedged language and anthropomorphic expressions."],"supporting_citations":[{"why":"One of the two recent surveys from which the 40 LLM UQ method papers were collected for analysis.","marker":"Shorinwa et al., 2024"},{"why":"The second survey used as a source of LLM UQ method papers and their evaluation benchmarks.","marker":"Huang et al., 2024"},{"why":"Provides the four criteria C1–C4 that the paper adopts for scoring the ecological validity of benchmarks.","marker":"Chaleshtori et al., 2024"},{"why":"Supplies the concept of ecological validity and the list of ways NLP datasets typically lack it.","marker":"De Vries et al., 2020"},{"why":"WildChat data showing only 6.3% of real user queries are factual QA, used to argue the benchmark task mix is unrepresentative.","marker":"Zhao et al., 2024"},{"why":"Human study showing uncalibrated algorithmic advice can outperform calibrated advice, supporting the claim that calibration is not sufficient for human uplift.","marker":"Vodrahalli et al., 2022"},{"why":"Theoretical result that calibrated models are suboptimal when aggregating human and algorithmic predictions, supporting the metrics critique.","marker":"Corvelo Benz and Rodriguez, 2023"},{"why":"Gives a necessary and sufficient condition for human-algorithm complementarity, used to argue that calibration alone does not guarantee uplift.","marker":"Donahue et al., 2022"},{"why":"Proposes FolkTexts, a dataset with intentional aleatoric uncertainty, used as an example for recommendation R4 and evidence that methods can fail under aleatoric uncertainty.","marker":"Cruz et al., 2024"}],"fun_headline_variants":["Better calibration scores don't yield better LLM collaboration","Survey of 40 methods: UQ benchmarks ignore human utility","LLM uncertainty metrics are out of touch with real users","To help users, LLM uncertainty must leave the benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's quantitative evidence depends on the authors' subjective annotations in Table 1 being correct (for example, that Natural Questions passes C1 while TriviaQA does not) and on the 40 sampled methods being representative of the broader LLM UQ literature.","fun_headline_variants_meta":{"raw":{"variants":["Better calibration scores don't yield better LLM collaboration","Survey of 40 methods: UQ benchmarks ignore human utility","LLM uncertainty metrics are out of touch with real users","To help users, LLM uncertainty must leave the benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000321,"raw_usage":{"total_tokens":1767,"prompt_tokens":864,"completion_tokens":903,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":836}},"tokens_in":480,"tokens_out":903,"duration_ms":10991,"temperature":1.0,"reasoning_tokens":836,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:33:19.103823+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human-subject study in which presenting calibrated LLM confidence scores to users produces a measurable, consistent improvement in decision accuracy on a realistic task, while uncalibrated scores do not, would directly contradict the claim that calibration metrics are unconnected to downstream utility. Alternatively, an independent re-annotation of the 22 benchmarks that placed the majority in different C1 or C2 categories would weaken the ecological-validity counts.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"WildChat data showing only 6.3% of real user queries are factual QA, used to argue the benchmark task mix is unrepresentative."}],"review_version":1}