{"id":"c5d6c840-dabe-4c23-a5c2-6c301203073e","arxiv_id":"2412.20892","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Uncertainty, predictive performance, and data dispersion can be cleanly separated through a decision-theoretic lens, and BALD is best seen as an estimator of short-run parameter information gain rather than a direct measure of reducible predictive uncertainty.","lead":"This paper argues that the common split of uncertainty into aleatoric and epistemic parts is incoherent, and replaces it with a decision-theoretic framework that derives uncertainty from the expected loss of a Bayes-optimal action. A reader should care because the framework clarifies what different uncertainty measures actually mean, and it explains when and why the popular BALD score is a good or bad guide for choosing new data.","discovery_kind":"paradigm_shift","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No internal inconsistency; the only load-bearing assumption is the stipulated loss/decision problem, which is explicitly acknowledged and does not invalidate the conditional central claim.","rationale":"The reader's weakest assumption identifies the same condition I would flag: the framework presupposes a specified decision problem and loss. I agree this is the most delicate point. But I do not treat it as a load-bearing objection because the paper never claims to deliver a loss-free, universal uncertainty measure. It explicitly starts from the decision of interest, acknowledges that specifying ℓ is hard in practice, and embraces the consequence that uncertainty measures are decision-maker-relative. The internal logic from that starting point is sound: Bayes-optimal actions, minimal expected loss, expected uncertainty reduction, and the irreducible/reducible decomposition all follow from the stated definitions. The BALD discussion is also carefully scoped with assumptions and explicit disclaimers, and the core figures are backed by available code. There is no circular step or hidden unsupported premise that would change the verdict. The only useful additional check is to probe how robust the short-run BALD interpretation is under misspecification, but even if that interpretation degrades, the main conceptual contribution of the paper stands.","tokens_in":17829,"tokens_out":13297,"duration_ms":138574,"concrete_test":"Re-run the Figure 4 protocol with a misspecified model (e.g., Student-t data modeled as Gaussian, or a heteroscedastic likelihood modeled homoscedastically) and compare εθ and εz as functions of n. This tests whether the claim that BALD is better understood as an estimator of short-run parameter information gain is robust outside the exactly well-specified conjugate settings used in the paper. If εθ is not smaller than εz in misspecified regimes, the practical BALD interpretation is limited, though the conceptual decision-theoretic framework is unaffected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central construction is a conditional theorem: given a loss ℓ and an expected-utility decision rule, h[p_n(z)] is the minimal expected loss and the §5.2 decomposition follows mathematically. The soft point is exactly loss specification: if there is no final decision (e.g., exploratory analysis, OOD detection, model reporting), the framework does not single out one uncertainty measure, and choosing ℓ reintroduces some of the arbitrariness the paper aims to remove. This is not an internal inconsistency: §3.2 and §5.1 explicitly acknowledge that ℓ can be hard to choose and that different decision-makers will use different measures. It does limit the scope of the replacement claim to settings with specified preferences, but the central conceptual contribution and the BALD analysis do not depend on providing a loss-free measure. The BALD propositions are framed as Bayes estimators under stated assumptions, and §5.5 explicitly disclaims generality, so the empirical demonstrations support rather than overclaim. The minimax alternative mentioned in §3.3 also means the expected-utility decision rule is itself a choice, but the paper presents it as one principled framework rather than the only possible one.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper argues that the standard aleatoric/epistemic uncertainty framework in machine learning is conceptually overloaded and internally incoherent, and proposes replacing it with a decision-theoretic perspective. The central construction defines predictive uncertainty h[p_n(z)] as the subjective expected loss of the Bayes-optimal action under the predictive model, and then derives an expected uncertainty reduction (EUR) that depends explicitly on the data-generating process and the data-acquisition policy. The paper also relates this uncertainty measure to predictive performance and data dispersion through proper scoring rules and discrepancy functions, and revisits the BALD score, arguing that it is best viewed as an imperfect estimator of short-run parameter information gain rather than a direct measure of long-run reducible predictive uncertainty. The claims are supported by formal propositions, illustrative examples, and exact conjugate-model experiments.","tokens_in":18046,"tokens_out":13552,"duration_ms":127550,"significance":"If the framework holds up, this is a valuable conceptual contribution: it unifies variance, entropy, and other uncertainty measures under a single decision-theoretic umbrella, provides a rigorous decomposition of reducible and irreducible uncertainty, and clarifies the status of BALD as an estimator rather than a target quantity. The paper is unusually honest about its limitations, explicitly acknowledging that the framework requires a specified loss and decision problem, and that the practical estimators rely on approximations. The empirical demonstrations in Figure 4 use exact conjugate models where the claimed estimation errors can be computed without approximation, which is a strength. The main weakness is a technical gap in the formal justification of Propositions 2, 3, and 6, which misapply a standard Bayes-estimator result to target quantities that are not functions of the model parameter under the stated assumptions.","major_comments":[{"comment":"The proofs of Propositions 2, 3, and 6 apply Proposition 1 to a quantity F that is not, under the stated assumptions, a function of the model parameter θ. For Proposition 2, F = E_{p_eval(z)}[s(p_n,z)] while f(θ) = E_{p_n(z|θ)}[s(p_n,z)]; the assumption that p_n(z) is 'a model intended to directly approximate' p_eval(z) does not imply F = f(θ), yet Proposition 1 requires the quantity of interest to be a pushforward of p_n(θ) through f. The argument therefore establishes only that h[p_n(z)] is the posterior mean of f(θ), not a Bayes estimator of the external expected score. The same gap appears in Proposition 3, where F = h[p_eval(z)] and f(θ) = h[p_n(z|θ)], and in Proposition 6, where F = E_{ptrain(z)}[H[p_n(θ|z)]] and f(θ) = E_{p_n(z|θ)}[H[p_n(θ|z)]]. The propositions are salvageable by adding a well-specifiedness assumption, for example p_eval(z) = p_n(z|θ_0) for some θ_0 (and similarly for ptrain in Proposition 6), but as written they are not proved. Because Section 5.4 uses these propositions to reinterpret entropy-based quantities as Bayes estimators, this is a load-bearing gap that should be fixed before publication.","section":"Section 5.4, Propositions 2 and 3 (and Proposition 6, Appendix B)"},{"comment":"The proof of Proposition 4 is too terse for a formal proposition. The statement 'reasoning about θ is equivalent to reasoning about y+1:∞' requires a careful argument that, under the stated posterior-consistency assumption, H[p_n(z|y+1:∞)] = H[p_n(z|θ∞)] almost surely and that this convergence is strong enough to justify applying Proposition 1 to the limit quantity. As written, the proof skips these steps and instead cites Fong et al. (2023) without spelling out the correspondence. This is not necessarily an error, but it needs to be expanded so that the reader can verify that the Bayes-estimator claim is not presupposing the desired conclusion.","section":"Appendix B, Proposition 4"}],"minor_comments":[{"comment":"The proof writes F = E_{ptrain(z)}[H[p_n(θ|z_{1:m})]], but the quantity being estimated, EIG^true_θ, involves a single new observation z = y_{n+1}; the subscript z_{1:m} appears to be a typo and should be z.","section":"Appendix B, Proposition 6 proof"},{"comment":"The convention that updating stochasticity is 'implicitly absorbed into y_{1:n}' is confusing, since y_{1:n} was earlier defined as training data. Consider introducing an explicit random seed variable and treating the machine-learning method as a deterministic function of data and seed.","section":"Section 3.4"},{"comment":"The caption and text use slightly different notation for the sample sizes: the text says n ∈ (1, 10, 100, 1000) and the caption says n ∈ {1, 10, 100, 1000}; please standardize.","section":"Section 5.5, Figure 4"},{"comment":"The proof would be clearer if it showed the intermediate step E_{p_n(θ)}[E_{p_n(z|θ)}[s(p_n,z)]] = E_{p_n(z)}[s(p_n,z)] = h[p_n(z)], since this is the key algebraic step connecting f(θ) to the predictive uncertainty.","section":"Section 5.4, Proposition 2 proof"}],"recommendation":"major_revision","confidential_remarks":"The paper makes a strong conceptual contribution and the central decision-theoretic framework is sound. The main concern is the proof gap in Propositions 2, 3, and 6, which is localized and fixable by restating the assumptions as well-specified models. If the authors address this, I would be supportive of acceptance. I am not concerned about the paper's strong claim that the aleatoric-epistemic view should be replaced, since the paper is appropriately explicit about the need for a specified loss and decision problem."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. The paper does two things well. First, it builds uncertainty measures from a specified decision problem and loss, so variance, entropy, and proper scoring rules all fall out of one construction (h[p] = expected loss of Bayes-optimal action). Second, it reinterprets BALD as a Bayes estimator of two distinct targets: the infinite-step predictive information gain and the one-step true parameter information gain, and shows empirically that the latter is the better match. The unification is real; Dawid, DeGroot, and Berger supply the parts, but the assembly and the BALD analysis are new. The paper ships code for the core figures, uses exact conjugate models so there are no hidden approximations, and explicitly disclaims generality where it should.\n\nThe soft spots are in proportion. The central construction is conditional on a loss function that reflects the decision-maker's preferences. For exploratory analysis, OOD detection, or model reporting, there often is no such final decision, and the framework does not single out a unique measure; the choice of ℓ reintroduces some of the arbitrariness the paper aims to remove. The paper acknowledges this in §3.2 and §5.1, which counts in its favor, but the abstract's 'in place of the aleatoric-epistemic view' is stronger than the conditional result delivers. The minimax alternative in §3.3 is also a reminder that expected-utility is itself a principled choice, not the only one.\n\nThe Bayes-estimator propositions are the thinnest part formally. They rely on posterior consistency and, in Prop 6, the construction of f(θ) is terse enough that I would want a referee to spell out the model-correctness assumption behind it. But the claims are framed carefully: BALD is an estimator, not a direct measure, and the figure shows the estimation error clearly. The citation pattern is fine; the self-citations are for the Figure 5 experimental setups, which is appropriate.\n\nWho is this for? Anyone who uses 'aleatoric' and 'epistemic' language operationally, and especially active-learning researchers who rely on BALD. It deserves a serious referee; this is the kind of paper that shapes how a subfield talks. My recommendation: send it to review, and ask the authors to tighten the framing so the replacement claim is clearly scoped to settings with a specified decision problem.","headline":"A genuinely useful decision-theoretic reframing of uncertainty in ML; the BALD reinterpretation is the sharpest new result, and the paper is honest about the cost of needing a specified decision problem.","tokens_in":18552,"tokens_out":3520,"would_cite":true,"duration_ms":33636,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62C05","62F15","68T37"],"pacs":[],"model":"deepseek-v4-flash","headline":"Given a decision and a loss, a rigorous uncertainty measure follows; BALD is an estimator of short-run parameter information gain, not a direct measure of reducible predictive uncertainty.","keywords":["aleatoric uncertainty","epistemic uncertainty","decision theory","subjective expected loss","expected uncertainty reduction","BALD","proper scoring rules","Bayesian active learning"],"falsifier":"Run the paper's conjugate-model experiment with misspecified generative models, such as data drawn from a distribution outside the assumed Beta-Bernoulli or Normal-Gamma family, and measure whether BALD still estimates short-run expected parameter information gain better than long-run predictive information gain; if misspecification reverses that ordering, the paper's explanation of BALD's utility is falsified. Separately, exhibiting a real deployment with no specified decision or loss yet one unambiguously agreed uncertainty measure would weaken the loss-grounding claim.","tokens_in":17677,"feed_emoji":"🎲","tokens_out":10053,"duration_ms":92469,"temperature":0.7,"pith_summary":"The paper claims that the standard aleatoric/epistemic view of uncertainty in machine learning is incoherent and too coarse: researchers attach many different mathematical quantities to each of the two concepts, blurring the line between what a model believes, how data are generated, and how a model should be evaluated. The proposed alternative starts from a final decision with an explicit loss: predictive uncertainty is the subjective expected loss of the Bayes-optimal action, $h[p_n(z)] = \\mathbb{E}_{p_n(z)}[\\ell(a^*_n,z)]$, which generalizes variance and Shannon entropy as two special cases. Reducibility of uncertainty then requires specifying the data-generating process, giving a well-defined expected uncertainty reduction and a decomposition into irreducible and reducible parts. The paper also separates predictive uncertainty from predictive performance and data dispersion, and reinterprets BALD as an estimator of short-run expected parameter information gain rather than a direct measure of long-run reducible predictive uncertainty. This matters because active-learning objectives and uncertainty-based methods are currently designed and evaluated on conceptual foundations the paper argues are shaky.","feed_headline":"BALD is an estimator, not a direct measure of reducible uncertainty","feed_subtitle":"Choose your loss function, and the right uncertainty measure follows; predictive performance and data dispersion stay separate.","key_machinery":"The load-bearing object is the minimal subjective expected loss $h[p_n(z)] = \\mathbb{E}_{p_n(z)}[\\ell(a^*_n,z)]$, which the paper calls predictive uncertainty; it is what makes the choice of uncertainty measure a derived quantity rather than a modeling preference. Two identities carry the rest of the argument: the expected uncertainty reduction $EUR^{\\mathrm{true}}_z(\\pi,m)$ based on the actual data-generating process, which yields the irreducible-reducible decomposition, and the discrepancy function $d(p_n,p_{\\mathrm{eval}}) = \\mathbb{E}_{p_{\\mathrm{eval}}(z)}[s(p_n,z)] - h[p_{\\mathrm{eval}}(z)]$, which separates uncertainty, predictive performance, and data dispersion while generalizing classical decompositions.","core_discovery":"On the paper's own terms, the central discovery is that measures of predictive uncertainty need not be chosen arbitrarily: once a decision problem $(A,Z,\\ell)$ is fixed, the minimal subjective expected loss of the Bayes-optimal action, $h[p_n(z)] = \\mathbb{E}_{p_n(z)}[\\ell(a^*_n,z)]$, is the uncertainty measure, with squared error yielding variance and log loss yielding entropy. From there the paper defines expected uncertainty reduction under an explicit data-acquisition policy, $EUR^{\\mathrm{true}}_z(\\pi,m) = \\mathbb{E}_{p_{\\mathrm{train}}(y^+_{1:m}\\mid\\pi)}[h[p_n(z)] - h[p_{n+m}(z)]]$, and shows the infinite-data limit gives a rigorous irreducible-reducible decomposition that applies to any method mapping data to a predictive distribution. A discrepancy identity, $d(p_n,p_{\\mathrm{eval}}) = \\mathbb{E}_{p_{\\mathrm{eval}}(z)}[s(p_n,z)] - h[p_{\\mathrm{eval}}(z)]$, separates predictive performance from uncertainty and data dispersion, generalizing bias-variance and cross-entropy/KL decompositions. Finally, the paper argues that BALD, the expected information gain in model parameters, is best understood as an estimator of the true one-step expected parameter information gain, not a direct measure of long-run reducible predictive uncertainty; the two can diverge substantially at finite $n$, which the paper demonstrates with conjugate models.","pith_inferences":["If the loss-based derivation is taken seriously, then in prediction-oriented active learning the acquisition objective should be the expected reduction in predictive uncertainty under the final loss, which is what prediction-oriented information gain objectives approximate; this would make the paper's view a direct recipe for choosing acquisition functions rather than only a reinterpretation of ex","The framework implies that the phrase 'epistemic uncertainty' is doing the work of at least three distinct quantities: parameter uncertainty, reducible predictive uncertainty, and short-run expected parameter information gain, so future empirical comparisons that treat 'epistemic uncertainty' as one number will struggle to be meaningful.","A concrete extension the paper does not run is, for a fixed decision problem with a specified loss, comparing acquisition functions derived from finite-m expected uncertainty reduction under model-simulated data against BALD and predictive entropy on the same benchmarks; if the derived functions do not match or beat loss-appropriate baselines, the practical relevance of the decision-theoretic deri"],"forward_implications":["Two decision-makers with different loss functions are not disagreeing about the same object when one reports variance and the other entropy; each is reporting the uncertainty measure their decision problem requires.","A reducible/irreducible decomposition is available for any data-to-predictive-distribution method, provided the data-generating policy is stated; stochastic parameters and exact Bayesian updating are not required.","BALD's practical success in active learning is compatible with it being a poor estimator of long-run predictive information gain, because acquisition horizons are short and it tracks one-step parameter information gain better.","Uncertainty alone gives no reliable signal of predictive performance or data dispersion; externally grounded evaluation is required."],"supporting_citations":[{"why":"Supplies the core result that minimal expected loss under the Bayes-optimal action measures uncertainty, and provides the discrepancy function used to link uncertainty, performance, and dispersion.","marker":"Dawid, 1998"},{"why":"Early decision-theoretic treatment of uncertainty and sequential experiments that grounds the expected-loss measure of uncertainty.","marker":"DeGroot, 1962"},{"why":"Provides the decision-theoretic information gain objective that the paper's expected uncertainty reduction generalizes.","marker":"Neiswanger et al, 2022"},{"why":"Introduced the BALD score as expected information gain in model parameters, the quantity the paper reinterprets as a short-run estimator.","marker":"Houlsby et al, 2011"},{"why":"Established the additive aleatoric-plus-epistemic decomposition and made BALD a standard active-learning acquisition objective.","marker":"Gal et al, 2017"},{"why":"The main target of the incoherence analysis: the paper shows this work attaches multiple distinct mathematical quantities to each of the two uncertainty concepts.","marker":"Kendall & Gal, 2017"},{"why":"Provides the prediction-oriented active learning framework whose empirical comparisons with BALD and predictive entropy appear in the paper's Figure 5.","marker":"Bickford Smith et al, 2023"},{"why":"Supplies further active-learning results and the predictive information gain objective, used to support the claim that prediction-oriented acquisition can beat BALD.","marker":"Bickford Smith et al, 2024"},{"why":"Convergence results under which the model's predictive distribution recovers the data-generating process, used in the propositions that characterize BALD as a Bayes estimator.","marker":"Doob, 1949; Freedman, 1963; 1965"}],"fun_headline_variants":["Uncertainty measure depends on your loss function","BALD estimates parameter info gain, not reducible uncertainty","Separating predictive performance, uncertainty, and dispersion","A decision-theoretic take on aleatoric and epistemic uncertainty","Your loss function defines your measure of uncertainty"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The derivation assumes the user can specify a final decision problem with a loss function that reflects their preferences; for many machine-learning deployments no such loss exists, and without one the framework does not single out a unique uncertainty measure.","fun_headline_variants_meta":{"raw":{"variants":["Uncertainty measure depends on your loss function","BALD estimates parameter info gain, not reducible uncertainty","Separating predictive performance, uncertainty, and dispersion","A decision-theoretic take on aleatoric and epistemic uncertainty","Your loss function defines your measure of uncertainty"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000377,"raw_usage":{"total_tokens":2013,"prompt_tokens":959,"completion_tokens":1054,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":978}},"tokens_in":575,"tokens_out":1054,"duration_ms":9716,"temperature":1.0,"reasoning_tokens":978,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:07:27.616971+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's conjugate-model experiment with misspecified generative models, such as data drawn from a distribution outside the assumed Beta-Bernoulli or Normal-Gamma family, and measure whether BALD still estimates short-run expected parameter information gain better than long-run predictive information gain; if misspecification reverses that ordering, the paper's explanation of BALD's utility is falsified. Separately, exhibiting a real deployment with no specified decision or loss yet one unambiguously agreed uncertainty measure would weaken the loss-grounding claim.","supporting_citations":[{"cited_title":"Coherent measures of discrepancy, uncertainty and dependence, with applications to Bayesian predictive experimental design","cited_arxiv_id":null,"evidence_quote":"Supplies the core result that minimal expected loss under the Bayes-optimal action measures uncertainty, and provides the discrepancy function used to link uncertainty, performance, and dispersion."},{"cited_title":"Uncertainty, information, and sequential experiments","cited_arxiv_id":null,"evidence_quote":"Early decision-theoretic treatment of uncertainty and sequential experiments that grounds the expected-loss measure of uncertainty."},{"cited_title":"Generalizing Bayesian optimization with decision-theoretic entropies","cited_arxiv_id":null,"evidence_quote":"Provides the decision-theoretic information gain objective that the paper's expected uncertainty reduction generalizes."},{"cited_title":"Bayesian active learning for classification and preference learning","cited_arxiv_id":null,"evidence_quote":"Introduced the BALD score as expected information gain in model parameters, the quantity the paper reinterprets as a short-run estimator."},{"cited_title":"What uncertainties do we need in Bayesian deep learning for computer vision? Conference on Neural Information Processing Systems","cited_arxiv_id":null,"evidence_quote":"The main target of the incoherence analysis: the paper shows this work attaches multiple distinct mathematical quantities to each of the two uncertainty concepts."},{"cited_title":"Prediction-oriented Bayesian active learning","cited_arxiv_id":null,"evidence_quote":"Provides the prediction-oriented active learning framework whose empirical comparisons with BALD and predictive entropy appear in the paper's Figure 5."},{"cited_title":"Making better use of unlabelled data in Bayesian active learning","cited_arxiv_id":null,"evidence_quote":"Supplies further active-learning results and the predictive information gain objective, used to support the claim that prediction-oriented acquisition can beat BALD."},{"cited_title":"Application of the theory of martingales","cited_arxiv_id":null,"evidence_quote":"Convergence results under which the model's predictive distribution recovers the data-generating process, used in the propositions that characterize BALD as a Bayes estimator."}],"review_version":1}