REVIEW 4 major objections 5 minor 50 references
On the Need to Align Intent and Implementation in Uncertainty Quantification for Machine Learning
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A paper argues that most ML uncertainty estimates in science are used for claims they were never designed to support, a failure it names 'construct drift'.
desk verdict A sincere organizational framework for UQ alignment; the checklist is useful, but the core diagnostic criterion needs a sharper definition before it becomes a standard. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the epistemic contract: a three-part alignment among the estimation target (prediction, parameter inference, indirect inference, simulator-parameter inference, unique-event forecast), the warrant that legitimizes uncertainty for that target, and the uncertainty construct (frequentist, Bayesian, fiducial, logical) that expresses it. The paper also supplies a practical testing apparatus: three axes of trustworthiness (formal guarantees, empirical reliability, model correspondence) and a scientific simulation-based inference checklist (theory check, forward checks, inverse checks, degeneracy mapping, global structure comprehension) that operationalize the contract.
What would settle it
Run the paper's own diagnostics on a set of method-first simulation-based inference papers: check whether each reported posterior passes simulation-based calibration, a posterior predictive check, and a sensitivity analysis to simulator perturbation; if most reported uncertainties pass all three, the claim that construct drift is widespread collapses, while if few pass, the diagnosis is confirmed.
Extended reading notes
Core claim
The paper's central claim is that the validity of an uncertainty estimate is not a property of the method alone but of the method paired with its inferential target and context. A variance across deep-ensemble members is a dispersion of model outputs, not by itself a measure of epistemic ignorance, and a prediction interval is a statement about future observables, not about latent physical parameters. When such quantities are invoked for the other role, the paper says the estimate suffers construct drift and trans-semantic transfer: the surface form of a guarantee is preserved while its justificatory grounding is lost. The remedy is an explicit epistemic contract that declares the estimation target, the warrant (long-run coverage, belief coherence, error control, or evidential support), and the construct that carries that warrant, then tests the result along formal, empirical, and domain-correspondence axes.
Load-bearing premise
The argument assumes that estimation targets are separable enough that each has a single, defensible uncertainty construct, so that reusing a construct across targets can be called drift; if hybrid targets or constructs that legitimately transfer across semantics are common, the diagnosis loses its force.
Editorial extensions
If this is right
- Every uncertainty report should declare its inference chain: target, decision goal, construct, and warrant, so that readers can see what the number is allowed to mean.
- Prediction intervals, credible regions, and ensemble variances are not interchangeable; using one in another's role voids the guarantee it carries.
- Simulation-based inference pipelines should run both forward checks (posterior predictive) and inverse checks (simulation-based calibration), and should test sensitivity to simulator perturbations.
- Evaluation should be engineered to the decision, for example stratified calibration when false-negative rates or phase boundaries are what matter.
- Borrowed terms such as 'epistemic', 'systematic', and 'confidence' need to be defined with respect to both their statistical and scientific context.
Reading between the lines
- This framework points to a concrete governance tool the paper leaves implicit: an 'uncertainty card' or construct declaration attached to published results, making target-warrant alignment auditable.
- Because Table 1 marks simulator-parameter inference under frequentist and fiducial warrants as under-explored, a natural next step is developing and calibrating uncertainty constructs for neural posterior and ratio estimators under model misspecification, not just under simulation.
- The three-axes checklist is also a lens for method comparison: two methods with equal coverage could be ranked by model correspondence, which current benchmark culture largely ignores.
- If construct drift is as widespread as the paper asserts, one testable consequence is that re-analyzing published simulation-based inference results with explicit target-construct alignment will change some scientific conclusions; this could be checked on a corpus.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that uncertainty quantification in scientific machine learning suffers from 'construct drift': statistical quantities are computed for one object (e.g., a prediction interval) and then invoked to support conclusions about another (e.g., latent physical parameters), without a defensible epistemic link. The paper proposes a taxonomy of six estimation targets (Section 2), four families of uncertainty constructs (Section 3), a mapping between them in Table 2, and three 'axes of trustworthiness' (formal guarantees, empirical reliability, model correspondence) in Section 5. Section 6 presents illustrative cases of misalignment, and Section 7 offers practical recommendations including declaring the inference chain, forward and inverse checks, and simulator-based stress tests. The contribution is explicitly organizational and conceptual rather than methodological.
Significance. If made precise, the framework could give scientific ML researchers a useful vocabulary for diagnosing when an uncertainty estimate does not support the claim it is used for. The paper usefully synthesizes ideas from statistics, philosophy of probability, and current SBI practice, and its practical recommendations—especially forward and inverse validation, simulator stress tests, and explicit inference-chain declarations—are valuable and actionable. The paper also correctly identifies real dangers in trans-semantic transfers, such as using prediction intervals as parameter constraints. However, the absence of an operational criterion for 'epistemic justification' currently limits the framework's force, and several empirical claims about prevalence are not supported. The central diagnosis is plausible and worth publishing, but the manuscript needs substantial revision to make the diagnostic criterion precise and to qualify or evidence its stronger assertions.
major comments (4)
- [Sections 1 and 5] The paper's central concept, 'epistemic justification', is never defined operationally. Section 1 calls it 'a defensible link' between the reported quantity and the claim it supports, but no criterion is given for when such a link is defensible. Section 5 introduces the three axes as 'necessary (if not sufficient)' for trustworthy UQ, yet Section 6 repeatedly infers construct drift from a violated axis (e.g., 'Variance versus uncertainty (violates Axes 1 & 2)'). Failing a non-sufficient check is not evidence of invalidity. The authors need to provide an operational criterion—for instance, a decision-theoretic condition, a coherence condition, or a formal link between the estimand and the reported quantity—so that 'construct drift' can be applied and tested.
- [Section 6, 'Variance versus uncertainty'] The flagship example of deep-ensemble variance as epistemic uncertainty is not decisive as stated. The authors claim that ensemble variance 'carries no formal calibration guarantee and is rarely tested for frequentist coverage or Bayesian coherence' and therefore lacks epistemic justification. This conflates absence of a particular guarantee with absence of justification. Under a Bayesian model-averaging interpretation, an ensemble can approximate a posterior over weights, and the variance across ensemble members can represent reducible epistemic uncertainty without needing to be a calibrated interval. The authors should either acknowledge this legitimate reading and restrict their claim to settings where no such interpretation holds, or supply a criterion that distinguishes legitimate Bayesian use from the alleged drift.
- [Section 6, 'Most SBI papers do not check this'] The paper states that 'Most SBI papers do not check this' regarding posterior predictive checks and sensitivity to simulator parameters. This is an empirical prevalence claim, but no survey, corpus analysis, or quantitative evidence is provided. Because the paper's motivation rests on the failure mode being widespread, this assertion is load-bearing. The authors should either provide systematic evidence (e.g., an analysis of a sample of SBI publications) or explicitly characterize the claim as anecdotal and temper its role in the argument.
- [Section 2 and Table 2] The taxonomy assumes that estimation targets form a small, separable set and that each target has one defensible construct pairing, as shown in Table 2. Many scientific workflows involve hybrid targets or constructs that are legitimately transferred across semantic levels—for example, using a posterior predictive distribution for experimental design, or using a prediction interval as an approximate constraint on a parameter in an embedded model. The paper does not discuss how to distinguish such legitimate transfers from construct drift. Without a treatment of hybrid cases, Table 2's 'defensible' pairings are asserted rather than derived, and the diagnostic loses its force in precisely the ambiguous cases where it is most needed.
minor comments (5)
- [Abstract and Section 2] The abstract contains the typo 'a illustrative suite'; Section 2, item 6 reads 'Simulation–based inference: : Likelihood-Free Learning' with a double colon. These should be corrected.
- [References] References [23] and [45] are the same paper (Talts et al., 'Validating Bayesian inference algorithms with simulation-based calibration') cited twice with different numbers; the duplicate should be removed or unified.
- [Section 7] The phrase 'epistemic hygiene--–' in the final paragraph contains stray hyphens and dashes; it should read 'epistemic hygiene'.
- [Section 6, 'Variance versus uncertainty'] The claim that deep ensembles are 'often used' to report variance as uncertainty in physics is not accompanied by a specific citation. Adding concrete references or explicitly labeling the statement as an illustrative pattern would strengthen the discussion.
- [Table 1] The entry characterizing frequentist SBI as 'under-explored' is debatable, since there is a body of work on frequentist coverage properties of simulation-based estimators, calibrated ABC, and confidence distributions. The table should either cite that literature or qualify the entry as an assessment of principled frequentist constructs specifically, not of all frequentist work in SBI.
Circularity Check
No significant circularity: the paper is an explicitly organizational position piece; its core taxonomy is stipulated, not derived, and the two self-citations are illustrative only.
full rationale
The paper performs no derivations, fits no parameters, and makes no quantitative predictions; its central notion of construct drift is introduced by stipulative definition ('a variance over model outputs is treated as a stand-in for model ignorance... We call this phenomenon construct drift') rather than by a derivation from inputs that would reduce to itself. Table 2 pairings are explicitly illustrative ('The figure is not exhaustive, but it signals which kinds of construct–target pairings are defensible'), so they are not presented as first-principles results. Self-citations [43,44] appear once as examples of conformal prediction losing subgroup explainability and carry no load-bearing justificatory weight for the framework. The unsupported empirical claim that 'Most SBI papers do not check this' is a correctness and evidence concern, not circularity. Accordingly, no step in the paper's argument reduces by construction to its own inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption The taxonomy of estimation targets (estimation, prediction, inference, predictive inference, indirect inference, SBI) is exhaustive for scientific ML workflows.
- domain assumption Each uncertainty construct family carries a single, non-transferable epistemic warrant; using it for another target is invalid.
- domain assumption The three trustworthiness axes (formal guarantees, empirical reliability, model correspondence) are necessary conditions for trustworthy UQ.
Cite this review
Pith. "Pith review of On the Need to Align Intent and Implementation in Uncertainty Quantification for Machine Learning." pith.science (2026). https://pith.science/paper/RC54I6U3
@misc{pith2026250603037,
author = {Pith},
title = {Pith review of: On the Need to Align Intent and Implementation in Uncertainty Quantification for Machine Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/RC54I6U3}},
note = {Machine review of arXiv:2506.03037}
}
read the original abstract
Quantifying uncertainties for machine learning (ML) models is a foundational challenge in modern data analysis. This challenge is compounded by at least two key aspects of the field: (a) inconsistent terminology surrounding uncertainty and estimation across disciplines, and (b) the varying technical requirements for establishing trustworthy uncertainties in diverse problem contexts. In this position paper, we aim to clarify the depth of these challenges by identifying these inconsistencies and articulating how different contexts impose distinct epistemic demands. We examine the current landscape of estimation targets (e.g., prediction, inference, simulation-based inference), uncertainty constructs (e.g., frequentist, Bayesian, fiducial), and the approaches used to map between them. Drawing on the literature, we highlight and explain examples of problematic mappings. To help address these issues, we advocate for standards that promote alignment between the \textit{intent} and \textit{implementation} of uncertainty quantification (UQ) approaches. We discuss several axes of trustworthiness that are necessary (if not sufficient) for reliable UQ in ML models, and show how these axes can inform the design and evaluation of uncertainty-aware ML systems. Our practical recommendations focus on scientific ML, offering illustrative cases and use scenarios, particularly in the context of simulation-based inference (SBI).
Reference graph
Works this paper leans on
-
[1]
Wiley Publications in Statistics, 1954
Leonard Savage.The Foundations of Statistics. Wiley Publications in Statistics, 1954
work page 1954
-
[2]
José M Bernardo and Adrian FM Smith.Bayesian theory, volume 405. John Wiley & Sons, 2009
work page 2009
-
[3]
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery. Strictly proper scoring rules, prediction, and estimation. Journal of the American statistical Association, 102(477):359–378, 2007
2007
-
[4]
Rajendra Acharya, Vladimir Makarenkov, and Saeid Nahavandi
Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, Dana Rezazadegan, Li Liu, Mohammad Ghavamzadeh, Paul Fieguth, Xiaochun Cao, Abbas Khosravi, U. Rajendra Acharya, Vladimir Makarenkov, and Saeid Nahavandi. A review of uncertainty quantification in deep learning: Techniques, applications and challenges.Information Fusion, 76:243–297, December 2021. ISSN 1566-2...
-
[5]
Simple and scalable predictive uncertainty estimation using deep ensembles, 2017
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles, 2017. URL https://arxiv.org/ abs/1612.01474
arXiv 2017
-
[6]
Weight uncertainty in neural networks, 2015
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra. Weight uncertainty in neural networks, 2015. URLhttps://arxiv.org/abs/1505.05424
arXiv 2015
-
[7]
Springer, 2005
Vladimir V ovk, Alexander Gammerman, and Glenn Shafer.Algorithmic learning in a random world, volume 29. Springer, 2005
2005
-
[8]
The frontier of simulation-based inference
Kyle Cranmer, Johann Brehmer, and Gilles Louppe. The frontier of simulation-based inference. Proceedings of the National Academy of Sciences, 117(48):30055–30062, 2020
2020
Show all 50 references
-
[9]
Science and statistics.Journal of the American Statistical Association, 71 (356):791–799, 1976
George EP Box. Science and statistics.Journal of the American Statistical Association, 71 (356):791–799, 1976
1976
-
[10]
Oberkampf and Christopher J
William L. Oberkampf and Christopher J. Roy.Verification and Validation in Scientific Comput- ing. Cambridge University Press, 2010
2010
-
[11]
Cambridge university press, 2014
Shai Shalev-Shwartz and Shai Ben-David.Understanding machine learning: From theory to algorithms. Cambridge university press, 2014
2014
-
[12]
Statistical inference.Australia: Duxbury/Thomson Learning, 2002
Casella G Berger RL. Statistical inference.Australia: Duxbury/Thomson Learning, 2002
2002
-
[13]
Bayesian data analysis, 3rd edn london, 2013
Andrew Gelman, John B Carlin, Hal S Stern, David B Dunson, Aki Vehtari, and Donald B Rubin. Bayesian data analysis, 3rd edn london, 2013
2013
-
[14]
Conformalized quantile regression
Yaniv Romano, Evan Patterson, and Emmanuel Candes. Conformalized quantile regression. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, ed- itors,Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. UR...
2019
-
[15]
The limits of distribution-free conditional predictive inference.Information and Inference: A Journal of the IMA, 10(2):455–482, 2021
Rina Foygel Barber, Emmanuel J Candes, Aaditya Ramdas, and Ryan J Tibshirani. The limits of distribution-free conditional predictive inference.Information and Inference: A Journal of the IMA, 10(2):455–482, 2021
2021
-
[16]
Indirect inference.Journal of applied econometrics, 8(S1):S85–S118, 1993
Christian Gourieroux, Alain Monfort, and Eric Renault. Indirect inference.Journal of applied econometrics, 8(S1):S85–S118, 1993
1993
-
[17]
Springer, 1977
John Wilder Tukey et al.Exploratory data analysis, volume 2. Springer, 1977
1977
-
[18]
Springer, 2005
Olav Kallenberg et al.Probabilistic symmetries and invariance principles, volume 9. Springer, 2005. 10
2005
-
[19]
Fastϵ -free inference of simulation models with bayesian conditional density estimation
George Papamakarios and Iain Murray. Fastϵ -free inference of simulation models with bayesian conditional density estimation. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors,Advances in Neural Information Processing Systems, volume 29. Curran Associates, ...
2016
-
[20]
Sequential neural likelihood: Fast likelihood-free inference with autoregressive flows
George Papamakarios, David Sterratt, and Iain Murray. Sequential neural likelihood: Fast likelihood-free inference with autoregressive flows. In Kamalika Chaudhuri and Masashi Sugiyama, editors,Proceedings of the Twenty-Second International Conference on Artifi- cial Intellige...
2019
-
[21]
Likelihood-free mcmc with amortized approximate ratio estimators, 2020
Joeri Hermans, V olodimir Begy, and Gilles Louppe. Likelihood-free mcmc with amortized approximate ratio estimators, 2020. URLhttps://arxiv.org/abs/1903.04057
2020 arXiv
-
[22]
Transmission of Justification and Warrant
Luca Moretti and Tommaso Piazza. Transmission of Justification and Warrant. In Edward N. Zalta and Uri Nodelman, editors,The Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University, Summer 2023 edition, 2023
2023
-
[24]
Cambridge University Press, 2006
Ian Hacking.The emergence of probability: A philosophical study of early ideas about probability, induction and statistical inference. Cambridge University Press, 2006
2006
-
[25]
Outline of a theory of statistical estimation based on the classical theory of probability.Philosophical Transactions of the Royal Society of London
Jerzy Neyman. Outline of a theory of statistical estimation based on the classical theory of probability.Philosophical Transactions of the Royal Society of London. Series A, Mathematical and Physical Sciences, 236(767):333–380, 1937
1937
-
[26]
University of Chicago Press, 1996
Deborah G Mayo.Error and the growth of experimental knowledge. University of Chicago Press, 1996
1996
-
[27]
David Roxbee Cox.Planning of experiments.Wiley, 1958
1958
-
[28]
John Wiley & Sons, 2017
Bruno De Finetti.Theory of probability: A critical introductory treatment. John Wiley & Sons, 2017
2017
-
[29]
OUP Oxford, 2004
Luc Bovens and Stephan Hartmann.Bayesian epistemology. OUP Oxford, 2004
2004
-
[30]
On fiducial inference.The Annals of Mathematical Statistics, 32(3):661–676, 1961
Donald AS Fraser. On fiducial inference.The Annals of Mathematical Statistics, 32(3):661–676, 1961
1961
-
[31]
Fiducial inference, then and now
Philip Dawid. Fiducial inference, then and now. InHandbook of Bayesian, Fiducial, and Frequentist Inference, pages 83–105. Chapman and Hall/CRC, 2024
2024
-
[32]
Generalized fiducial inference: A review and new results.Journal of the American Statistical Association, 111(515):1346–1361, 2016
Jan Hannig, Hari Iyer, Randy CS Lai, and Thomas CM Lee. Generalized fiducial inference: A review and new results.Journal of the American Statistical Association, 111(515):1346–1361, 2016
2016
-
[33]
Courier Corporation, 2013
John Maynard Keynes.A treatise on probability. Courier Corporation, 2013
2013
-
[34]
Citeseer, 1962
Rudolf Carnap.Logical foundations of probability, volume 2. Citeseer, 1962
1962
-
[35]
Cambridge university press, 2003
Edwin T Jaynes.Probability theory: The logic of science. Cambridge university press, 2003
2003
-
[36]
Verification, validation, and predictive capability in computational engineering and physics.Appl
William L Oberkampf, Timothy G Trucano, and Charles Hirsch. Verification, validation, and predictive capability in computational engineering and physics.Appl. Mech. Rev., 57(5): 345–384, 2004
2004
-
[37]
Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods.Machine learning, 110(3):457–506, 2021
Eyke Hüllermeier and Willem Waegeman. Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods.Machine learning, 110(3):457–506, 2021. 11
2021
-
[38]
Aleatory or epistemic? does it matter?Structural safety, 31(2):105–112, 2009
Armen Der Kiureghian and Ove Ditlevsen. Aleatory or epistemic? does it matter?Structural safety, 31(2):105–112, 2009
2009
-
[39]
Explainable uncertainty quantifications for deep learning-based molecular property prediction.Journal of Cheminformatics, 15(1):13, 2023
Chu-I Yang and Yi-Pei Li. Explainable uncertainty quantifications for deep learning-based molecular property prediction.Journal of Cheminformatics, 15(1):13, 2023
2023
-
[40]
Bayesian astrostatistics: a backward look to the future
Thomas J Loredo. Bayesian astrostatistics: a backward look to the future. InAstrostatistical challenges for the new astronomy, pages 15–40. Springer, 2012
2012
-
[41]
Towards reliable simulation-based inference with balanced neural ratio estimation.Advances in Neural Information Processing Systems, 35:20025–20037, 2022
Arnaud Delaunoy, Joeri Hermans, François Rozet, Antoine Wehenkel, and Gilles Louppe. Towards reliable simulation-based inference with balanced neural ratio estimation.Advances in Neural Information Processing Systems, 35:20025–20037, 2022
2022
-
[42]
Bayes and frequentism: a particle physicist’s perspective.Contemporary Physics, 54(1):1–16, February 2013
Louis Lyons. Bayes and frequentism: a particle physicist’s perspective.Contemporary Physics, 54(1):1–16, February 2013. ISSN 1366-5812. doi: 10.1080/00107514.2012.756312. URL http://dx.doi.org/10.1080/00107514.2012.756312
2013
-
[43]
Conformal prediction with temporal quantile adjustments.Advances in Neural Information Processing Systems, 35:31017–31030, 2022
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. Conformal prediction with temporal quantile adjustments.Advances in Neural Information Processing Systems, 35:31017–31030, 2022
2022
-
[44]
Conformal prediction intervals with temporal dependence.arXiv preprint arXiv:2205.12940, 2022
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. Conformal prediction intervals with temporal dependence.arXiv preprint arXiv:2205.12940, 2022
2022 arXiv
-
[45]
Validating bayesian inference algorithms with simulation-based calibration, 2020
Sean Talts, Michael Betancourt, Daniel Simpson, Aki Vehtari, and Andrew Gelman. Validating bayesian inference algorithms with simulation-based calibration, 2020. URL https://arxiv. org/abs/1804.06788
2020 arXiv
-
[46]
Bayesianly justifiable and relevant frequency calculations for the applied statistician.The Annals of Statistics, pages 1151–1172, 1984
Donald B Rubin. Bayesianly justifiable and relevant frequency calculations for the applied statistician.The Annals of Statistics, pages 1151–1172, 1984
1984
-
[47]
Posterior predictive assessment of model fitness via realized discrepancies.Statistica sinica, pages 733–760, 1996
Andrew Gelman, Xiao-Li Meng, and Hal Stern. Posterior predictive assessment of model fitness via realized discrepancies.Statistica sinica, pages 733–760, 1996
1996
-
[48]
Predicting good probabilities with supervised learning
Alexandru Niculescu-Mizil and Rich Caruana. Predicting good probabilities with supervised learning. InProceedings of the 22nd international conference on Machine learning, pages 625–632, 2005
2005
-
[49]
The comparison and evaluation of forecasters
Morris H DeGroot and Stephen E Fienberg. The comparison and evaluation of forecasters. Journal of the Royal Statistical Society: Series D (The Statistician), 32(1-2):12–22, 1983
1983
-
[50]
Cherian, and Emmanuel J
Isaac Gibbs, John J. Cherian, and Emmanuel J. Candès. Conformal prediction with conditional guarantees, 2024. URLhttps://arxiv.org/abs/2305.12616
2024 arXiv
-
[51]
Uncertainty cards,
Biao Wu, Haihui Zhang, Yuanxun Zhou, Lanting Zhang, and Hong Wang. Uncertainty quantification of phase boundary in a composition-phase map via bayesian strategies.Phys. Rev. Mater., 7:025201, Feb 2023. doi: 10.1103/PhysRevMaterials.7.025201. URL https: //link.aps.org/doi/10.11...
2023 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.