Pith. sign in

REVIEW 1 major objections 5 minor 213 references

Current LLMs are consistent but miscalibrated when turning probabilistic predictions into words, especially for uncertainty, so they are not yet reliable standalone risk communicators.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 07:01 UTC pith:GK72M7VI

load-bearing objection Clean factorial study showing LLMs are consistent but miscalibrated verbalizers of likelihood/uncertainty, with the bottleneck in verbalization itself; metric is load-bearing but well-ablated. the 1 major comments →

arxiv 2607.03882 v2 pith:GK72M7VI submitted 2026-07-04 cs.CL cs.AI

Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language

classification cs.CL cs.AI
keywords large language modelsrisk communicationcalibrationconsistencynatural language explanationspredictive uncertaintylikelihood verbalizationpost-hoc explanation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Large language models are being asked to turn numerical predictions from other AI systems into plain-language risk statements that people can use. For that role to work, the same numbers must get the same words every time, and the chosen words must match the true size of the risk and the true amount of uncertainty. This paper tests nine models on exactly that job. It feeds them samples drawn from Beta distributions that systematically vary likelihood (mode) and uncertainty (prior sample size), places the samples in six different real-world contexts, and forces each model to pick from ordered lists of standard verbal descriptors. The models are mostly consistent across repeated identical queries, yet they systematically pick descriptors that do not track the numerical magnitudes, with far worse performance for uncertainty than for likelihood. Handing the models the true mode and sample size up front reduces sensitivity to context but does not fix the miscalibration. The authors conclude that the failure is in the verbalization step itself and that current LLMs are not yet safe zero-shot tools for standalone probabilistic risk communication.

Core claim

Nine LLMs, tested across systematically varied likelihood and uncertainty, six domain contexts, ten temperatures, and ten repetitions, are generally consistent yet miscalibrated when forced to select verbal descriptors for probabilistic predictions; performance is substantially weaker for uncertainty than for likelihood, and supplying precomputed mode and prior sample size reduces context sensitivity without curing the miscalibration, locating the bottleneck in the verbalization step rather than in statistical inference.

What carries the argument

A controlled two-stage pipeline that draws N=100 samples from Beta distributions parameterized by mode m and prior sample size k, inserts them into domain-specific prompts, forces selection from ordered seven-term likelihood and uncertainty descriptor lists, and scores the outputs with a consistency measure (identical descriptors across ten identical-input repetitions) plus a modified Jonckheere-Terpstra calibration score that counts order violations of the model's own usage and applies a vocabulary-coverage penalty.

Load-bearing premise

The claim that a forced choice from fixed seven-term lists scored by a modified Jonckheere-Terpstra statistic with a vocabulary-coverage penalty correctly measures whether an LLM can communicate risk in natural language.

What would settle it

Re-run the identical grid of (m,k) Beta samples and domain contexts but allow free-form generation (or a validated open vocabulary of hedges) and re-score with human expert or lay probability mappings; if the same models then produce calibrated, ordered verbalizations, the paper's conclusion that the verbalization bottleneck is fundamental would be overturned.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. The paper evaluates whether nine LLMs can serve as reliable zero-shot post-hoc explainers of probabilistic predictions by selecting verbal descriptors for likelihood and uncertainty. Upstream predictions are simulated as N=100 samples from Beta distributions on a full factorial grid of modes m and prior sample sizes k; models are prompted under six domain contexts, ten temperatures, and ten repetitions, with guided decoding to a fixed seven-term Spiegelhalter-style NLE list. Consistency is measured by agreement across identical-prompt repetitions; calibration is measured by a modified Jonckheere-Terpstra trend test (with vocabulary-coverage penalty) that checks whether descriptor choices respect the numerical ordering of m (likelihood) or k (uncertainty). Results show high consistency but substantial miscalibration, especially for uncertainty; providing precomputed m and k reduces context sensitivity without curing miscalibration, locating the bottleneck in verbalization. Six metric-robustness ablations (external expert midpoints, β-sweep, 5-descriptor scale, empirical moments, variance/entropy alternatives, stake/scope decomposition) preserve the qualitative pattern.

Significance. If the result holds, it supplies a controlled, multi-model, multi-context demonstration that current LLMs are not yet dependable standalone risk-communication tools for probabilistic AI outputs, with the failure concentrated on uncertainty verbalization rather than numerical inference. The design is unusually thorough for the area: full factorial m/k grid, temperature sweep, domain framing, guided decoding, pass@k filtering, 95% CIs, public code, and six ablations that stress-test the calibration metric itself. GPT-5.4’s near-perfect likelihood mapping further shows the metric can register success when present. The work therefore gives a concrete, falsifiable baseline for future verbalization interventions and for claims about LLMs as post-hoc explainers in high-stakes settings.

major comments (1)
  1. The central claim is supported by the experimental design and the ablations. The only load-bearing operationalization that still warrants explicit discussion is the modified JT score with β=0.5 coverage penalty and forced seven-term lists (§4.4.2, Appendix A.5.2). The paper already runs the necessary stress tests (external Spiegelhalter/IPCC midpoints, β-sweep including β=0, 5-descriptor scale, empirical moments, variance/entropy/CI-width alternatives); all preserve the qualitative pattern. In revision, the authors should state more prominently in the main text (not only in the ablation subsection) that the “no architecture above 0.40” claim is specific to concentration-based scoring and that variance-based scoring lifts top models above that threshold, so readers do not over-generalize the absolute magnitude.
minor comments (5)
  1. Figure 2 and Appendix A.7: confidence intervals are often invisible at the plotted scale; consider reporting numerical CIs in a companion table or using error bars that remain legible.
  2. Table 1 and §4.2: the uncertainty descriptor list is presented as ordered from least to most certain, but the verbal labels mix “certain” and “uncertain” poles; a short note on polarity would help readers parse the JT groups.
  3. §5.5 and Discussion: the stake/scope decomposition is under-powered (missing medium-population cell); flag this more clearly as a design limitation rather than only as a non-finding.
  4. Appendix A.6 and front-page link: ensure the anonymized GitHub is replaced by the permanent public repository before camera-ready.
  5. Typos and polish: “degreetowhich” (§2), occasional missing spaces around citations, and “Gastronomy” vs “restaurant recommendations” inconsistency between Table 2 and the domain list.

Circularity Check

0 steps flagged

No significant circularity: pure empirical evaluation with ablated free parameters and external benchmarks; no derivation reduces to its inputs by construction.

full rationale

The paper is a controlled empirical study of nine LLMs on forced-choice verbalization of likelihood and uncertainty from synthetic Beta samples. Consistency is measured by agreement across identical-prompt repetitions; calibration is a modified Jonckheere–Terpstra score that checks whether a model’s own descriptor ordering respects the numerical order of modes (m) or prior sample sizes (k), plus an explicit vocabulary-coverage penalty at β=0.5. Neither metric is derived from a first-principles claim that collapses into its inputs: the ground-truth m/k grid is known by construction of the simulation, the descriptor lists are taken from Spiegelhalter et al. (2011), and the precomputed-m/k ablation, external expert-midpoint rescoring (Spiegelhalter/IPCC), β-sweep (including β=0), 5-descriptor ablation, empirical-sample vs population-parameter scoring, and variance/entropy/CI-width alternatives all leave the qualitative pattern intact. Free parameters are stated and stress-tested rather than fitted and re-presented as predictions. Related-work citations (including one co-author paper on self-explanations) are background, not load-bearing uniqueness theorems. There is therefore no self-definitional loop, no fitted-input-called-prediction, and no self-citation chain that forces the central claim. Score 0 is the correct honest finding.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 1 invented entities

Empirical evaluation paper. Load-bearing choices are the Beta generative model, the fixed NLE vocabularies drawn from Spiegelhalter/IPCC, the modified JT calibration metric with its coverage penalty, and the assumption that forced single-token selection under guided decoding fairly tests zero-shot risk communication. No new physical entities are postulated; free parameters are experimental knobs that are ablated.

free parameters (4)
  • β (vocabulary-coverage penalty exponent) = 0.5
    Set to 0.5 for the main power-penalty calibration score; controls how heavily models are penalized for unused descriptors. Sensitivity sweep is reported but the primary numbers use this fixed value.
  • m and k grids (mode and prior sample size) = m ∈ {0.05…0.95}, k ∈ {50…950}
    Ten evenly spaced medians of uniform bins over [0,1] and [0,1000]; define the 100 Beta distributions that serve as ground-truth likelihood/uncertainty pairs.
  • N = 100 samples per distribution = 100
    Number of Monte-Carlo-style draws presented to the LLM; chosen to emulate uncertainty-aware upstream models.
  • temperature grid = 0.0, 0.1, …, 0.9
    Ten values 0.0–0.9 used to probe sampling stochasticity; not fitted but free experimental design choice.
axioms (4)
  • domain assumption Beta distributions parameterized by mode m and prior sample size k adequately simulate the probabilistic outputs of an upstream uncertainty-aware model (e.g., Monte-Carlo Dropout).
    Stated in §4.2; all ground-truth likelihood and uncertainty values derive from this generative model.
  • domain assumption The ordered seven-term likelihood and uncertainty NLE lists (from Spiegelhalter et al. 2011 / IPCC) form a valid ordinal scale for calibration scoring.
    Table 1 and §4.2; the JT statistic treats these lists as strictly ordered groups.
  • ad hoc to paper A model’s own past descriptor usage defines the correct numerical ordering for calibration (self-consistency of ordering).
    Core of the modified JT metric in §4.4.2; external expert midpoints are used only in a robustness ablation.
  • ad hoc to paper Forced selection of a single descriptor under guided decoding is a fair test of zero-shot risk-communication ability.
    §4.2–4.3; unconstrained generation is acknowledged as future work but not tested.
invented entities (1)
  • Modified Jonckheere-Terpstra calibration score with logarithmic/power vocabulary-coverage penalty no independent evidence
    purpose: Quantifies how often a model’s NLE choices violate the ordering implied by its own usage, while penalizing incomplete vocabulary use.
    Defined in §4.4.2 and Appendix A.5.2; central dependent variable. Independent evidence is limited to the external-expert ablation that largely corroborates the ranking.

pith-pipeline@v1.1.0-grok45 · 20539 in / 3004 out tokens · 28087 ms · 2026-07-13T07:01:34.368173+00:00 · methodology

0 comments
read the original abstract

LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural language. For this role to be viable, models must produce identical verbal descriptions for identical inputs, and select descriptions that accurately reflect the magnitude of the underlying numerical quantities. We evaluate whether nine LLMs meet these requirements within a two-stage prediction pipeline, in which an upstream model has produced probabilistic outputs characterized by their likelihood and uncertainty, and LLMs are tasked with selecting an appropriate verbal descriptor for each. We simulate predictions from an upstream model by taking samples from a Beta distribution parameterized by its mode and prior sample size. We then prompt LLMs to explain these predictions under six domain contexts and with ten temperature settings, and repeating each experiment ten times. We find that LLMs are generally consistent but miscalibrated, with substantially weaker performance on uncertainty than on likelihood tasks. Providing models with precomputed summary statistics (mode and prior sample size) reduced sensitivity to contextual framing but did not resolve the underlying miscalibration, suggesting that the bottleneck resides in the verbalization step itself. These findings indicate that current LLMs do not yet constitute reliable zero-shot standalone risk communication tools for probabilistic predictions.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

213 extracted references · 38 linked inside Pith

  1. [1]

    Attention is All you Need , url =

    Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, ukasz and Polosukhin, Illia , booktitle =. Attention is All you Need , url =

  2. [2]

    Gerd Gigerenzer and Ralph Hertwig and Ewald H. H. Hoffrage and Ulrich Hoffrage , title =. Psychological Science in the Public Interest , year =

  3. [3]

    Artificial Intelligence Review , year =

    Hristos Tyralis and Georgia Papacharalampous , title =. Artificial Intelligence Review , year =

  4. [4]

    Murphy , title =

    Kevin P. Murphy , title =. 2022 , publisher =

  5. [5]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , year =

    Alex Kendall and Yarin Gal , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , year =

  6. [6]

    Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL) , year =

    Stephanie Lin and Jacob Hilton and Owain Evans , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL) , year =

  7. [7]

    International Journal of Statistics in Medical Research , year =

    Arif Ali and Abdur Rasheed and Afaq Ahmed Siddiqui and Maliha Naseer and Saba Wasim and Waseem Akhtar , title =. International Journal of Statistics in Medical Research , year =

  8. [8]

    2024 , eprint =

    Neil Band and Xuechen Li and Tengyu Ma and Tatsunori Hashimoto , title =. 2024 , eprint =

  9. [9]

    Teaching Language Models to Faithfully Express their Uncertainty , journal =

    Eikema, Bryan and Ilia, Evgenia and de Souza, Jos. Teaching Language Models to Faithfully Express their Uncertainty , journal =. 2025 , url =

  10. [10]

    2025 , url =

    Bijean Ghafouri and Shahrad Mohammadzadeh and James Zhou and Pratheeksha Nair and Jacob-Junqi Tian and Mayank Goel and Reihaneh Rabbany and Jean-François Godbout and Kellin Pelrine , title =. 2025 , url =

  11. [11]

    Jawad , title =

    Hashim, M. Jawad , title =. Ulster Medical Journal , volume =. 2024 , url =

  12. [12]

    Summary for Policymakers: Global Warming of 1.5°C , year =

  13. [13]

    Jackson and Katerina Andreadis and Jessica S

    Nicholas J. Jackson and Katerina Andreadis and Jessica S. Ancker , title =. JAMA Network Open , year =

  14. [14]

    2024 , eprint =

    Andreas Madsen and Sarath Chandar and Siva Reddy , title =. 2024 , eprint =

  15. [15]

    2021 , eprint =

    Matthias Minderer and Josip Djolonga and Rob Romijnders and Frances Hubis and Xiaohua Zhai and Neil Houlsby and Dustin Tran and Mario Lucic , title =. 2021 , eprint =

  16. [16]

    Science , year =

    David Spiegelhalter and Mike Pearson and Ian Short , title =. Science , year =

  17. [18]

    Consistency Score Metric Documentation , year =

  18. [19]

    arXiv.org , author=

    Explaining Chest X-ray Pathologies in Natural Language , url=. arXiv.org , author=

  19. [20]

    Bowman and Newton Cheng and Esin Durmus and Zac Hatfield-Dodds and Scott R

    Mrinank Sharma and Meg Tong and Tomasz Korbak and David Duvenaud and Amanda Askell and Samuel R. Bowman and Newton Cheng and Esin Durmus and Zac Hatfield-Dodds and Scott R. Johnston and Shauna Kravec and Timothy Maxwell and Sam McCandlish and Kamal Ndousse and Oliver Rausch and Nicholas Schiefer and Da Yan and Miranda Zhang and Ethan Perez , title =. 2023 , url =

  20. [22]

    2026 , eprint=

    A Survey of Large Language Models , author=. 2026 , eprint=

  21. [23]

    Annual Review of Statistics and Its Application , author=

    Risk and Uncertainty Communication , volume=. Annual Review of Statistics and Its Application , author=. 2017 , month=. doi:https://doi.org/10.1146/annurev-statistics-010814-020148 , number=

  22. [24]

    arXiv.org , author=

    Quantifying Uncertainty in Natural Language Explanations of Large Language Models , url=. arXiv.org , author=

  23. [25]

    2021 IEEE/CVF International Conference on Computer Vision (ICCV) , year=

    e-ViL: A Dataset and Benchmark for Natural Language Explanations in Vision-Language Tasks , author=. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) , year=

  24. [26]

    ML4CMH@AAAI , year=

    Natural Language Explanations for Suicide Risk Classification Using Large Language Models , author=. ML4CMH@AAAI , year=

  25. [27]

    Bioengineering , volume=

    Large Language Models in Healthcare and Medical Applications: A Review , author=. Bioengineering , volume=. 2025 , url=

  26. [28]

    Humanities and Social Sciences Communications , volume=

    Large Language Models in Legal Systems: A Survey , author=. Humanities and Social Sciences Communications , volume=. 2025 , url=

  27. [29]

    arXiv preprint arXiv:2406.11903 , year=

    A Survey of Large Language Models for Financial Applications: Progress, Prospects and Challenges , author=. arXiv preprint arXiv:2406.11903 , year=

  28. [30]

    Risk Analysis , author=

    Probability Information in Risk Communication: A Review of the Research Literature , volume=. Risk Analysis , author=. 2009 , month=. doi:https://doi.org/10.1111/j.1539-6924.2008.01137.x , number=

  29. [33]

    Proceedings of The 33rd International Conference on Machine Learning , pages =

    Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning , author =. Proceedings of The 33rd International Conference on Machine Learning , pages =. 2016 , editor =

  30. [34]

    Vllm.ai , author=

    Quickstart - vLLM , url=. Vllm.ai , author=

  31. [35]

    Advances in Large Margin Classifiers , author=

    Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods , volume=. Advances in Large Margin Classifiers , author=. 2000 , month=

  32. [36]

    arXiv.org , author=

    Obtaining Calibrated Probabilities from Boosting , url=. arXiv.org , author=

  33. [37]

    Prediction , url=

    Statistical Thinking - Classification vs. Prediction , url=. www.fharrell.com , author=. 2017 , month=

  34. [38]

    Nih.gov , author=

    Poor handling of continuous predictors in clinical prediction models using logistic regression: a systematic review , url=. Nih.gov , author=

  35. [41]

    arXiv.org , author=

    Scaling up Masked Diffusion Models on Text , url=. arXiv.org , author=

  36. [42]

    Anthropic.com , author=

    Labor market impacts of AI: A new measure and early evidence , url=. Anthropic.com , author=

  37. [44]

    Nature Machine Intelligence , author=

    The need for uncertainty quantification in machine-assisted medical decision making , volume=. Nature Machine Intelligence , author=. 2019 , month=. doi:https://doi.org/10.1038/s42256-018-0004-1 , number=

  38. [45]

    arXiv:1606.06565 [cs] , author=

    Concrete Problems in AI Safety , volume=. arXiv:1606.06565 [cs] , author=. 2016 , month=

  39. [46]

    A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions , volume=

    Huang, Lei and Yu, Weijiang and Ma, Weitao and Zhong, Weihong and Feng, Zhangyin and Wang, Haotian and Chen, Qianglong and Peng, Weihua and Feng, Xiaocheng and Qin, Bing and Liu, Ting , year=. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions , volume=. ACM transactions on office information systems ,...

  40. [47]

    and Nowozin, Sebastian and Dillon, Joshua V

    Ovadia, Yaniv and Fertig, Emily and Ren, Jie and Nado, Zachary and Sculley, D. and Nowozin, Sebastian and Dillon, Joshua V. and Lakshminarayanan, Balaji and Snoek, Jasper , year=. Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift , url=. doi:https://doi.org/10.48550/arXiv.1906.02530 , journal=

  41. [48]

    arXiv.org , author=

    Artificial Intelligence Index Report 2025 , url=. arXiv.org , author=

  42. [49]

    An Introduction to Statistical Learning , ISBN=

    James, Gareth and Witten, Daniela and Hastie, Trevor and Tibshirani, Robert , year=. An Introduction to Statistical Learning , ISBN=. doi:https://doi.org/10.1007/978-1-0716-1418-1 , journal=

  43. [50]

    Information Fusion , author=

    Explainable Artificial Intelligence (XAI): Concepts, taxonomies, Opportunities and Challenges toward Responsible AI , volume=. Information Fusion , author=. 2020 , month=. doi:https://doi.org/10.1016/j.inffus.2019.12.012 , number=

  44. [51]

    Why Should I Trust You?

    “Why Should I Trust You?”: Explaining the Predictions of Any Classifier , url=. arXiv.org , author=. 2016 , month=

  45. [52]

    arXiv:1705.07874 [cs, stat] , author=

    A Unified Approach to Interpreting Model Predictions , url=. arXiv:1705.07874 [cs, stat] , author=. 2017 , month=

  46. [53]

    arXiv:2201.11903 [cs] , author=

    Chain of Thought Prompting Elicits Reasoning in Large Language Models , url=. arXiv:2201.11903 [cs] , author=. 2022 , month=

  47. [54]

    arXiv.org , author=

    Counterfactual Explanations for Machine Learning: Challenges Revisited , url=. arXiv.org , author=

  48. [55]

    Explainability and uncertainty: Two sides of the same coin for enhancing the interpretability of deep learning models in healthcare , volume=

    Salvi, Massimo and Seoni, Silvia and Campagner, Andrea and Gertych, Arkadiusz and Acharya, U.Rajendra and Molinari, Filippo and Cabitza, Federico , year=. Explainability and uncertainty: Two sides of the same coin for enhancing the interpretability of deep learning models in healthcare , volume=. doi:https://doi.org/10.1016/j.ijmedinf.2025.105846 , journal=

  49. [56]

    Uncertainty-aware deep learning in healthcare: A scoping review , volume=

    Loftus, Tyler J and Shickel, Benjamin and Ruppert, Matthew M and Balch, Jeremy A and Tezcan Ozrazgat-Baslanti and Tighe, Patrick J and Efron, Philip A and Hogan, William R and Rashidi, Parisa and Upchurch, Gilbert R and Azra Bihorac , year=. Uncertainty-aware deep learning in healthcare: A scoping review , volume=. PLOS digital health , publisher=. doi:ht...

  50. [57]

    Medical Decision Making , author=

    Numeric, Verbal, and Visual Formats of Conveying Health Risks: Suggested Best Practices and Future Recommendations , volume=. Medical Decision Making , author=. 2007 , month=. doi:https://doi.org/10.1177/0272989x07307271 , number=

  51. [58]

    and Smyth, Padhraic , year=

    Steyvers, Mark and Tejeda, Heliodoro and Kumar, Aakriti and Belem, Catarina and Karny, Sheer and Hu, Xinyue and Mayer, Lukas W. and Smyth, Padhraic , year=. What large language models know and what people think they know , volume=. doi:https://doi.org/10.1038/s42256-024-00976-7 , journal=

  52. [59]

    Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning , url=

    Gal, Yarin and Ghahramani, Zoubin , year=. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning , url=. proceedings.mlr.press , publisher=

  53. [60]

    and van Loon, Kim and Kappen, Martinus A.M

    Kappen, Teus H. and van Loon, Kim and Kappen, Martinus A.M. and van Wolfswinkel, Leo and Vergouwe, Yvonne and van Klei, Wilton A. and Moons, Karel G.M. and Kalkman, Cor J. , year=. Barriers and facilitators perceived by physicians when using prediction models in practice , volume=. doi:https://doi.org/10.1016/j.jclinepi.2015.09.008 , journal=

  54. [61]

    Archives of Internal Medicine , author=

    The Effect of Giving Global Coronary Risk Information to Adults , volume=. Archives of Internal Medicine , author=. 2010 , month=. doi:https://doi.org/10.1001/archinternmed.2009.516 , number=

  55. [62]

    BMJ Open , author=

    Impact of provision of cardiovascular disease risk estimates to healthcare professionals and patients: a systematic review , volume=. BMJ Open , author=. 2015 , month=. doi:https://doi.org/10.1136/bmjopen-2015-008717 , number=

  56. [63]

    Statistics in Medicine , author=

    Dichotomizing continuous predictors in multiple regression: a bad idea , volume=. Statistics in Medicine , author=. 2005 , pages=. doi:https://doi.org/10.1002/sim.2331 , number=

  57. [64]

    BMC Medicine , author=

    Three myths about risk thresholds for prediction models , volume=. BMC Medicine , author=. 2019 , month=. doi:https://doi.org/10.1186/s12916-019-1425-3 , number=

  58. [65]

    Uncertainty of risk estimates from clinical prediction models: rationale, challenges, and approaches , url=

    Riley, Richard D and Collins, Gary S and Kirton, Laura and Snell, Kym Ie and Ensor, Joie and Whittle, Rebecca and Dhiman, Paula and Maarten van Smeden and Liu, Xiaoxuan and Alderman, Joseph and Krishnarajah Nirantharakumar and Manson-Whitton, Jay and Westwood, Andrew J and Cazier, Jean-Baptiste and Moons, Karel G M and Martin, Glen P and Sperrin, Matthew ...

  59. [66]

    Royal Society Open Science , author=

    Communicating uncertainty about facts, numbers and science , volume=. Royal Society Open Science , author=. 2019 , month=. doi:https://doi.org/10.1098/rsos.181870 , number=

  60. [67]

    arXiv:1612.01474 [cs, stat] , author=

    Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles , url=. arXiv:1612.01474 [cs, stat] , author=. 2017 , month=

  61. [68]

    The Well-Calibrated Bayesian , volume=

    Dawid, A P , year=. The Well-Calibrated Bayesian , volume=. Journal of the American Statistical Association , publisher=. doi:https://doi.org/10.2307/2287720 , number=

  62. [69]

    arXiv.org , author=

    A Survey of Confidence Estimation and Calibration in Large Language Models , url=. arXiv.org , author=

  63. [70]

    BMC Medicine , author=

    Calibration: the Achilles heel of predictive analytics , volume=. BMC Medicine , author=. 2019 , month=. doi:https://doi.org/10.1186/s12916-019-1466-7 , number=

  64. [71]

    arXiv.org , author=

    Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs , url=. arXiv.org , author=

  65. [72]

    Internet Archive , author=

    Notes on matters affecting the health, efficiency, and hospital administration of the British Army [electronic resource] : founded chiefly on the experience of the late war : Nightingale, Florence, 1820-1910 : Free Download, Borrow, and Streaming : Internet Archive , url=. Internet Archive , author=

  66. [73]

    DigitalCommons@RISD , author=

    Women and Work , url=. DigitalCommons@RISD , author=

  67. [74]

    Brewer and Julie S

    Baruch Fischhoff and Noel T. Brewer and Julie S. Downs , title =. 2011 , publisher =

  68. [75]

    Nature Climate Change , author=

    The role of social and decision sciences in communicating uncertain climate risks , volume=. Nature Climate Change , author=. 2011 , month=. doi:https://doi.org/10.1038/nclimate1080 , number=

  69. [76]

    Applied Cognitive Psychology , year =

    Kimihiko Yamagishi , title =. Applied Cognitive Psychology , year =

  70. [77]

    Science , year =

    Amos Tversky and Daniel Kahneman , title =. Science , year =

  71. [78]

    Levin and Sandra L

    Irwin P. Levin and Sandra L. Schneider and Gary J. Gaeth , title =. Organizational Behavior and Human Decision Processes , year =

  72. [79]

    Polack and Stephen J

    Fernando P. Polack and Stephen J. Thomas and Nicholas Kitchin and Judith Absalon and Alejandra Gurtman and Stephen Lockhart and John L. Perez and Gonzalo P\'. Safety and Efficacy of the. New England Journal of Medicine , year =. doi:10.1056/NEJMoa2034577 , url =

  73. [80]

    Social Science & Medicine , year =

    Stefania Pighin and Lucia Savadori and Silvia Bonalumi and Roberta Bonfiglioli and Constantinos Hadjichristidis , title =. Social Science & Medicine , year =

  74. [81]

    The Lancet Microbe , year =

    Piero Olliaro and Els Torreele and Michel Vaillant , title =. The Lancet Microbe , year =

  75. [82]

    Climatic Change , year =

    Shinichiro Asayama and Masahiro Sugiyama and Seita Emori and Fumiko Kasuga and Chieko Watanabe , title =. Climatic Change , year =

  76. [83]

    Kohl and Neil Stenhouse , title =

    Patrice A. Kohl and Neil Stenhouse , title =. Environmental Communication , year =

  77. [84]

    2025 , url =

    How Good Are. 2025 , url =

  78. [85]

    We have 12 years to limit climate change catastrophe, warns UN , url=

    Watts, Jonathan , year=. We have 12 years to limit climate change catastrophe, warns UN , url=. the Guardian , publisher=

  79. [86]

    One Earth , author=

    Now or Never: How Media Coverage of the IPCC Special Report on 1.5°C Shaped Climate-Action Deadlines , volume=. One Earth , author=. 2019 , month=. doi:https://doi.org/10.1016/j.oneear.2019.10.026 , number=

  80. [87]

    Oana-Maria Camburu and Tim Rockt. e-. Advances in Neural Information Processing Systems , volume =. 2018 , url =

Showing first 80 references.