{"id":"d4ab5305-a29b-4e7b-b974-6ab4dc527caf","arxiv_id":"2506.03037","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A position paper framing the mismatch between uncertainty quantities and their intended scientific claims as 'construct drift,' and proposing trustworthiness axes for scientific ML.","lead":"This position paper argues that uncertainty estimates in machine learning often support claims they were never designed to justify, such as reporting a variance over model outputs as if it were knowledge about the world. It proposes a diagnostic framework that matches uncertainty methods to the scientific question being asked, with practical advice for simulation-based inference.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim depends on an undefined notion of 'epistemic justification': no operational criterion separates construct drift from legitimate alternative interpretations, so a flagship example (deep-ensemble variance) can be justified under Bayesian model averaging, making the diagnosis…","rationale":"The paper is a position paper, not a method paper; its contribution is organizational. The taxonomy of targets (Section 2), the four construct families (Section 3), the trustworthiness axes (Section 5), and the SBI checklist (Section 5) are clearly presented and potentially useful for practitioners. My concern is not with the taxonomy per se but with the central claim's testability. The phrase 'lacks epistemic justification' is doing enormous work: it turns a normative preference for semantic tidiness into an alleged widespread empirical phenomenon. That move only works if 'epistemic justification' is defined precisely enough to classify cases. The paper never supplies such a definition, and its own axes are explicitly 'necessary (if not sufficient),' so violating an axis cannot establish that an uncertainty estimate is invalid for its claimed role. The deep-ensemble example illustrates the risk: under a Bayesian interpretation of ensembles, the reported variance is not a category error but an estimator of posterior predictive variance. If one accepted the paper's framing, a legitimate method could be falsely condemned. The concrete test I propose—constructing a Bayesian ensemble and checking whether the variance carries the epistemic meaning the paper denies—would settle whether the flagship example actually exhibits drift. If the example is justified, the paper needs a narrower, better-defined claim. The reader's verdict (CONDITIONAL) remains appropriate: the framework is valuable enough to warrant conditional acceptance, but the central empirical/diagnostic claim needs revision. I therefore do not propose changing the verdict, only strengthening the conditions: the authors should define 'epistemic justification' operationally and support the prevalence claim with a survey.","tokens_in":1076,"tokens_out":947,"duration_ms":107776,"concrete_test":"Take the Section 6 deep-ensemble example in a concrete Bayesian setting: specify a prior over network weights, approximate the posterior with an ensemble, and derive the posterior predictive variance. Show that the ensemble variance estimates the epistemic component and that the resulting predictive distribution passes simulation-based calibration and coverage checks on synthetic data. If this holds, the blanket assertion that such variance 'lacks epistemic justification' is false, and the paper must either sharpen its criterion or concede that construct drift is case-dependent. Alternatively, ask the authors to supply a formal decision procedure that classifies each row of Table 2 and the Section 6 examples into drift/non-drift; if the procedure is circular or absent, the central claim is not testable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's core diagnosis ('construct drift') requires that a reported uncertainty quantity have a determinate 'epistemic justification' relative to its intended role. Section 5's three axes are introduced as 'necessary (if not sufficient)' for trustworthy UQ, yet Section 6 repeatedly infers misalignment from a violated axis—e.g., 'Variance versus uncertainty (violates Axes 1 & 2)'—which does not follow: failing a non-sufficient check is not evidence of invalidity. More fundamentally, the paper never defines what makes a quantity-to-claim link 'defensible.' This matters because its flagship example, deep-ensemble variance reported as epistemic uncertainty, admits a rigorous Bayesian reading. If the ensemble approximates a posterior over weights, its predictive variance decomposes into aleatoric and epistemic components, and the epistemic component is exactly the reducible model uncertainty; it need not be a calibrated interval or a tested posterior to carry that meaning. The paper's objection—that such variance 'carries no formal calibration guarantee and is rarely tested for frequentist coverage or Bayesian coherence'—confuses absence of a particular guarantee with absence of justification. Without a precise criterion (e.g., a decision-theoretic or coherence condition), 'construct drift' is unfalsifiable: uses the authors approve of are 'aligned,' uses they disapprove of are 'drift.' Table 2 compounds this by listing two distinct warrants for simulator-based theta (NPE posterior and fiducial/confidence distribution), so the intended role is not unique. The framework's utility as a diagnostic therefore rests on an unstated and untested definition.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that uncertainty quantification in scientific machine learning suffers from 'construct drift': statistical quantities are computed for one object (e.g., a prediction interval) and then invoked to support conclusions about another (e.g., latent physical parameters), without a defensible epistemic link. The paper proposes a taxonomy of six estimation targets (Section 2), four families of uncertainty constructs (Section 3), a mapping between them in Table 2, and three 'axes of trustworthiness' (formal guarantees, empirical reliability, model correspondence) in Section 5. Section 6 presents illustrative cases of misalignment, and Section 7 offers practical recommendations including declaring the inference chain, forward and inverse checks, and simulator-based stress tests. The contribution is explicitly organizational and conceptual rather than methodological.","tokens_in":12507,"tokens_out":4092,"duration_ms":50735,"significance":"If made precise, the framework could give scientific ML researchers a useful vocabulary for diagnosing when an uncertainty estimate does not support the claim it is used for. The paper usefully synthesizes ideas from statistics, philosophy of probability, and current SBI practice, and its practical recommendations—especially forward and inverse validation, simulator stress tests, and explicit inference-chain declarations—are valuable and actionable. The paper also correctly identifies real dangers in trans-semantic transfers, such as using prediction intervals as parameter constraints. However, the absence of an operational criterion for 'epistemic justification' currently limits the framework's force, and several empirical claims about prevalence are not supported. The central diagnosis is plausible and worth publishing, but the manuscript needs substantial revision to make the diagnostic criterion precise and to qualify or evidence its stronger assertions.","major_comments":[{"comment":"The paper's central concept, 'epistemic justification', is never defined operationally. Section 1 calls it 'a defensible link' between the reported quantity and the claim it supports, but no criterion is given for when such a link is defensible. Section 5 introduces the three axes as 'necessary (if not sufficient)' for trustworthy UQ, yet Section 6 repeatedly infers construct drift from a violated axis (e.g., 'Variance versus uncertainty (violates Axes 1 & 2)'). Failing a non-sufficient check is not evidence of invalidity. The authors need to provide an operational criterion—for instance, a decision-theoretic condition, a coherence condition, or a formal link between the estimand and the reported quantity—so that 'construct drift' can be applied and tested.","section":"Sections 1 and 5"},{"comment":"The flagship example of deep-ensemble variance as epistemic uncertainty is not decisive as stated. The authors claim that ensemble variance 'carries no formal calibration guarantee and is rarely tested for frequentist coverage or Bayesian coherence' and therefore lacks epistemic justification. This conflates absence of a particular guarantee with absence of justification. Under a Bayesian model-averaging interpretation, an ensemble can approximate a posterior over weights, and the variance across ensemble members can represent reducible epistemic uncertainty without needing to be a calibrated interval. The authors should either acknowledge this legitimate reading and restrict their claim to settings where no such interpretation holds, or supply a criterion that distinguishes legitimate Bayesian use from the alleged drift.","section":"Section 6, 'Variance versus uncertainty'"},{"comment":"The paper states that 'Most SBI papers do not check this' regarding posterior predictive checks and sensitivity to simulator parameters. This is an empirical prevalence claim, but no survey, corpus analysis, or quantitative evidence is provided. Because the paper's motivation rests on the failure mode being widespread, this assertion is load-bearing. The authors should either provide systematic evidence (e.g., an analysis of a sample of SBI publications) or explicitly characterize the claim as anecdotal and temper its role in the argument.","section":"Section 6, 'Most SBI papers do not check this'"},{"comment":"The taxonomy assumes that estimation targets form a small, separable set and that each target has one defensible construct pairing, as shown in Table 2. Many scientific workflows involve hybrid targets or constructs that are legitimately transferred across semantic levels—for example, using a posterior predictive distribution for experimental design, or using a prediction interval as an approximate constraint on a parameter in an embedded model. The paper does not discuss how to distinguish such legitimate transfers from construct drift. Without a treatment of hybrid cases, Table 2's 'defensible' pairings are asserted rather than derived, and the diagnostic loses its force in precisely the ambiguous cases where it is most needed.","section":"Section 2 and Table 2"}],"minor_comments":[{"comment":"The abstract contains the typo 'a illustrative suite'; Section 2, item 6 reads 'Simulation–based inference: : Likelihood-Free Learning' with a double colon. These should be corrected.","section":"Abstract and Section 2"},{"comment":"References [23] and [45] are the same paper (Talts et al., 'Validating Bayesian inference algorithms with simulation-based calibration') cited twice with different numbers; the duplicate should be removed or unified.","section":"References"},{"comment":"The phrase 'epistemic hygiene--–' in the final paragraph contains stray hyphens and dashes; it should read 'epistemic hygiene'.","section":"Section 7"},{"comment":"The claim that deep ensembles are 'often used' to report variance as uncertainty in physics is not accompanied by a specific citation. Adding concrete references or explicitly labeling the statement as an illustrative pattern would strengthen the discussion.","section":"Section 6, 'Variance versus uncertainty'"},{"comment":"The entry characterizing frequentist SBI as 'under-explored' is debatable, since there is a body of work on frequentist coverage properties of simulation-based estimators, calibrated ABC, and confidence distributions. The table should either cite that literature or qualify the entry as an assessment of principled frequentist constructs specifically, not of all frequentist work in SBI.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a conceptual position piece; its normative claim is within scope for a journal that publishes such contributions, but the current draft overreaches in several places. The largest risk is that the central diagnosis is unfalsifiable without an operational criterion for epistemic justification. If the authors can supply such a criterion and temper the empirical claims, the paper could make a useful contribution. The journal should also consider whether a purely organizational contribution with no new method is within its scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honest summary: this is a sincere position paper that delivers an organizational framework, not a new algorithm or theorem. The genuinely useful pieces are the taxonomy of estimation targets, the four uncertainty-construct families, the three trustworthiness axes, and the SBI checklist in Section 5. The authors are upfront that the contribution is organizational, and the warnings—prediction intervals are not credible intervals; ensemble variance is not automatically epistemic uncertainty; conformal coverage does not buy causal insight—are real and worth saying in scientific ML.\n\nI also think the paper does several things well. It positions the critique relative to Box and Oberkampf/Roy rather than pretending the problem is new. It names a real failure mode in papers that report uncertainty without declaring the inferential target. The illustrative examples in Section 6, especially SED modeling and NRE with surrogates, are concrete. Self-citations [43,44] are only used as illustrations, so the citation pattern is not a problem.\n\nThe soft spots are the usual ones for a position paper, but they matter. The central notion—epistemic justification—is never operationalized. The three axes are introduced as necessary, not sufficient, so labeling a case as 'violates Axes 1 & 2' does not by itself prove the uncertainty is invalid. The deep-ensemble example is the clearest symptom: an ensemble can be read as approximate Bayesian model averaging, in which case the variance does carry an epistemic interpretation even without a calibration guarantee. The paper conflates absence of a particular guarantee with absence of any warrant. That is the weakest link. Also, 'Most SBI papers do not check this' is an empirical claim with no survey, and Table 1's 'under-explored' for frequentist SBI is debatable given the calibration literature.\n\nThese are not fatal. A position paper can be useful as a checklist and a vocabulary even when its central diagnostic criterion is informal. But the authors should be pushed to define what makes a quantity-to-claim link defensible, for example by giving a decision-theoretic or coherence condition, and to soften the prevalence claims or back them with a systematic review.\n\nI would bring this to a reading group and would cite it as a reference for 'declare the target and the construct.' It deserves a serious referee; with revision, it could become a standard reference for UQ evaluation in scientific ML.","headline":"A sincere organizational framework for UQ alignment; the checklist is useful, but the core diagnostic criterion needs a sharper definition before it becomes a standard.","tokens_in":13074,"tokens_out":2651,"would_cite":true,"duration_ms":31200,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A paper argues that most ML uncertainty estimates in science are used for claims they were never designed to support, a failure it names 'construct drift'.","keywords":["construct drift","epistemic warrant","uncertainty quantification","simulation-based inference","prediction intervals","trustworthy AI","Bayesian vs frequentist","scientific machine learning"],"falsifier":"Run the paper's own diagnostics on a set of method-first simulation-based inference papers: check whether each reported posterior passes simulation-based calibration, a posterior predictive check, and a sensitivity analysis to simulator perturbation; if most reported uncertainties pass all three, the claim that construct drift is widespread collapses, while if few pass, the diagnosis is confirmed.","tokens_in":12048,"feed_emoji":"🎯","tokens_out":5393,"duration_ms":60396,"temperature":0.7,"pith_summary":"This position paper argues that uncertainty estimates in machine learning, especially in scientific applications such as simulation-based inference, are frequently computed for one object and then used to support conclusions about another. It names this failure 'construct drift' and traces it to a missing epistemic contract: reported uncertainty has no defensible link among the estimation target, the warrant that justifies it, and the construct used to represent it. The paper builds a taxonomy of six estimation targets and four uncertainty-construct families, pairs them by warrant, and offers three axes for testing trustworthiness: formal guarantees, empirical reliability, and model correspondence. If the diagnosis holds, a large share of published uncertainty quantification in scientific ML is not valid for the claims it is used to support, and explicit intent-implementation alignment becomes a necessary standard.","feed_headline":"Why trust intervals in scientific ML are often the wrong intervals","feed_subtitle":"Prediction sets get treated as parameter evidence; the paper offers a target-warrant-contract fix.","key_machinery":"The load-bearing machinery is the epistemic contract: a three-part alignment among the estimation target (prediction, parameter inference, indirect inference, simulator-parameter inference, unique-event forecast), the warrant that legitimizes uncertainty for that target, and the uncertainty construct (frequentist, Bayesian, fiducial, logical) that expresses it. The paper also supplies a practical testing apparatus: three axes of trustworthiness (formal guarantees, empirical reliability, model correspondence) and a scientific simulation-based inference checklist (theory check, forward checks, inverse checks, degeneracy mapping, global structure comprehension) that operationalize the contract.","core_discovery":"The paper's central claim is that the validity of an uncertainty estimate is not a property of the method alone but of the method paired with its inferential target and context. A variance across deep-ensemble members is a dispersion of model outputs, not by itself a measure of epistemic ignorance, and a prediction interval is a statement about future observables, not about latent physical parameters. When such quantities are invoked for the other role, the paper says the estimate suffers construct drift and trans-semantic transfer: the surface form of a guarantee is preserved while its justificatory grounding is lost. The remedy is an explicit epistemic contract that declares the estimation target, the warrant (long-run coverage, belief coherence, error control, or evidential support), and the construct that carries that warrant, then tests the result along formal, empirical, and domain-correspondence axes.","pith_inferences":["This framework points to a concrete governance tool the paper leaves implicit: an 'uncertainty card' or construct declaration attached to published results, making target-warrant alignment auditable.","Because Table 1 marks simulator-parameter inference under frequentist and fiducial warrants as under-explored, a natural next step is developing and calibrating uncertainty constructs for neural posterior and ratio estimators under model misspecification, not just under simulation.","The three-axes checklist is also a lens for method comparison: two methods with equal coverage could be ranked by model correspondence, which current benchmark culture largely ignores.","If construct drift is as widespread as the paper asserts, one testable consequence is that re-analyzing published simulation-based inference results with explicit target-construct alignment will change some scientific conclusions; this could be checked on a corpus."],"forward_implications":["Every uncertainty report should declare its inference chain: target, decision goal, construct, and warrant, so that readers can see what the number is allowed to mean.","Prediction intervals, credible regions, and ensemble variances are not interchangeable; using one in another's role voids the guarantee it carries.","Simulation-based inference pipelines should run both forward checks (posterior predictive) and inverse checks (simulation-based calibration), and should test sensitivity to simulator perturbations.","Evaluation should be engineered to the decision, for example stratified calibration when false-negative rates or phase boundaries are what matter.","Borrowed terms such as 'epistemic', 'systematic', and 'confidence' need to be defined with respect to both their statistical and scientific context."],"supporting_citations":[{"why":"Introduces deep ensembles, the source of the practice where variance across model outputs is reported as epistemic uncertainty.","marker":"[5]"},{"why":"Supplies the conformal prediction framework, the exemplar of predictive inference with long-run coverage as its warrant.","marker":"[7]"},{"why":"Defines simulation-based inference and the contexts where learned posterior approximations are used for parameter inference.","marker":"[8]"},{"why":"Provides the caution about using statistical quantities outside their role, cited in the paper's definition of construct drift.","marker":"[9]"},{"why":"Establishes verification and validation distinctions that inform the paper's model-correspondence and trustworthiness axes.","marker":"[10]"},{"why":"Supplies the notion of epistemic warrant and its transmission, which grounds the paper's epistemic-contract framing.","marker":"[22]"},{"why":"Documents that neural ratio estimation can yield overconfident posteriors with poor frequentist coverage, a concrete case of misalignment.","marker":"[41]"},{"why":"Provides simulation-based calibration, the inverse-parameter-space diagnostic the paper recommends for validating posterior approximations.","marker":"[45]"}],"fun_headline_variants":["Why ML uncertainty often targets the wrong thing","Prediction intervals aren't evidence for parameters","Fix ML uncertainty with a target-warrant contract","Align intent and implementation in ML uncertainty","The epistemic contract fix for ML uncertainty"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that estimation targets are separable enough that each has a single, defensible uncertainty construct, so that reusing a construct across targets can be called drift; if hybrid targets or constructs that legitimately transfer across semantics are common, the diagnosis loses its force.","fun_headline_variants_meta":{"raw":{"variants":["Why ML uncertainty often targets the wrong thing","Prediction intervals aren't evidence for parameters","Fix ML uncertainty with a target-warrant contract","Align intent and implementation in ML uncertainty","The epistemic contract fix for ML uncertainty"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001322,"raw_usage":{"total_tokens":5377,"prompt_tokens":935,"completion_tokens":4442,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":4376}},"tokens_in":551,"tokens_out":4442,"duration_ms":33978,"temperature":1.0,"reasoning_tokens":4376,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:09:37.823976+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's own diagnostics on a set of method-first simulation-based inference papers: check whether each reported posterior passes simulation-based calibration, a posterior predictive check, and a sensitivity analysis to simulator perturbation; if most reported uncertainties pass all three, the claim that construct drift is widespread collapses, while if few pass, the diagnosis is confirmed.","supporting_citations":[{"cited_title":"Science and statistics.Journal of the American Statistical Association, 71 (356):791–799, 1976","cited_arxiv_id":null,"evidence_quote":"Provides the caution about using statistical quantities outside their role, cited in the paper's definition of construct drift."},{"cited_title":"Oberkampf and Christopher J","cited_arxiv_id":null,"evidence_quote":"Establishes verification and validation distinctions that inform the paper's model-correspondence and trustworthiness axes."},{"cited_title":"Transmission of Justification and Warrant","cited_arxiv_id":null,"evidence_quote":"Supplies the notion of epistemic warrant and its transmission, which grounds the paper's epistemic-contract framing."},{"cited_title":"Towards reliable simulation-based inference with balanced neural ratio estimation.Advances in Neural Information Processing Systems, 35:20025–20037, 2022","cited_arxiv_id":null,"evidence_quote":"Documents that neural ratio estimation can yield overconfident posteriors with poor frequentist coverage, a concrete case of misalignment."}],"review_version":1}