REVIEW 2 major objections 6 minor 24 references
Uncertainty quality tracks multi-modal diversity: methods that explore many loss basins set the standard, while shared-weight shortcuts and single-pass tricks trade that diversity away.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 13:22 UTC pith:GUBLO43X
load-bearing objection Solid scoped survey: method–measure split and mechanism×scope taxonomy are the real value; L/M/H tables are disclosed qualitative placements, not a hidden flaw. the 2 major comments →
Uncertainty quantification for trustworthy deep learning: Methods and measures
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Reading uncertainty quantification through a method–measure split and a generative-mechanism × network-scope taxonomy shows that between-mode, full-network exploration is what delivers top-tier uncertainty, while efficient shared-weight and single-basin methods systematically lose functional diversity and therefore weaker behavior under shift—yet carefully designed distance-aware single-pass models can still reach strong OOD detection without that multi-modal component.
What carries the argument
The method–measure separation plus the two-axis taxonomy (how the predictive ensemble is generated: explicit members, posterior sampling, or closed-form predictive; and network scope: full network, shared backbone/last layer, or single deterministic net). That grid, not architecture brand names, is what organizes the quality–cost trade-off and the comparison tables.
Load-bearing premise
That qualitative high/medium/low ratings stitched from different papers, architectures, datasets, and shifts—including some inferred only from method structure—are solid enough to rank whole method families head to head.
What would settle it
On a fixed architecture and the standard OpenOOD / shift suites, show that a cheap shared-weight or last-layer method matches a deep ensemble on both calibration under corruption severity and OOD AUROC once compute is matched; or show that last-layer diversity alone recovers most of the ensemble’s epistemic gap.
If this is right
- Practitioners should treat deep ensembles (and multi-modal extensions) as the quality reference and judge cheaper methods by how much functional diversity they keep under shift.
- Any sample-based ensemble can be paired with any sample-based uncertainty measure; closed-form single-pass methods only expose their built-in signal unless you sample them.
- Hybrid designs—small ensembles of distance-aware backbones, or one distance-aware trunk with several heads—become the natural next architecture class to test.
- Bounded pairwise scores and linear-time gated epistemic measures need the same standardized OOD/shift bake-off that methods already get.
- Generative language uncertainty must redefine diversity and the aleatoric/epistemic split over meanings, not fixed labels.
Where Pith is reading between the lines
- If diversity is the real scarce resource, evaluation suites should report effective ensemble size or representation disagreement under shift, not only ECE and AUROC.
- Safety cases that rely on single-pass deterministic uncertainty may be fine for far-OOD flagging yet still miss the multi-solution ignorance that full ensembles capture.
- The method–measure split suggests a modular software stack: one backend that emits member predictives, and swappable measure plugins—including conformal wrappers—without retraining.
- Epistemic collapse at large scale implies that “more parameters” alone will not buy trustworthy uncertainty; explicit multi-modal or repulsive mechanisms may stay necessary.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews uncertainty quantification for deep learning, scoped to ensemble-based and approximate Bayesian methods and to the measures that summarize their predictive ensembles. Its organizing contribution is a clean separation between the method that produces a predictive distribution and the measure that reads uncertainty from it, together with a two-axis taxonomy (generative mechanism × network scope; Table 2, Figure 1). Methods are grouped into BNNs, MC Dropout, deep ensembles, efficient ensemble approximations, and last-layer/single-pass approaches, with adjacent treatment of evidential/prior networks, conformal prediction, post-hoc calibration, OOD detection, selective prediction, and a brief LLM section. The authors consolidate evaluation practice (Section 11), diversity theory (Section 13), and entropy versus pairwise-divergence measures (Section 14), and close with open directions on last-layer diversity, efficient epistemic measures, shift, hybrids, and generative UQ.
Significance. If the organizational claims hold—as the loss-landscape and diversity literature largely supports—the paper offers a usable map of a crowded field that existing surveys treat more broadly or more architecture-first. Depth on efficient ensembles (BatchEnsemble, SWAG/MultiSWAG, CreDE, masksembles) and single-pass/last-layer methods (SNGP, DUQ/DDU/DUE, LLE, VBLL), plus an explicit method–measure decoupling and a consolidated evaluation basis for qualitative comparisons, are genuine contributions relative to Gawlikowski et al., Abdar et al., and He et al. The structural reading that full-network multi-modal exploration sets the quality standard while shared-weight/single-basin approximations lose functional diversity, with distance-aware single-pass methods as a partial OOD exception, is a clear and citable synthesis. Open problems (efficient non-entropy epistemic measures, last-layer diversity limits, hybrids, LLM semantic UQ) are well posed for follow-on work.
major comments (2)
- [Section 11–12, Table 3] Section 11–12 and Table 3 note (a)/†: the L/M/H OOD and calibration ratings are the main comparative device supporting family rankings, yet they pool heterogeneous architectures, datasets, and shift types, with several † ratings inferred from structure rather than direct benchmarks. The manuscript already discloses this and frames the cells as relative placements, which is appropriate for a survey; still, the gold-standard narrative and several M/H contrasts would be more robust if Table 3 (or a short appendix) marked which cells rest on OpenOOD/Uncertainty Baselines-style head-to-heads versus originating-paper or structural inference, so readers can weight the rankings accordingly.
- [Section 14.5–14.7, Tables 6–8] Section 14.5–14.7 and Tables 6–8: the structural case that EPJS removes EPKL’s unboundedness failure mode while avoiding MI’s mixture underestimation is carefully stated as non-empirical, and validation on standard OOD/shift suites is correctly listed as open (Section 17.2). Because the survey’s measure contribution partly rests on elevating pairwise and variance-gated alternatives over the default MI decomposition, a brief explicit caveat in Section 14.5 (and in Table 7’s EPJS/VGMU rows) that no claim of superior detection AUROC is being made would prevent over-reading by practitioners.
minor comments (6)
- [Section 14.7] Section 14.7 and Table 5/7: variance-gated ensembles, VGMU, and VGN are introduced largely via Gillis et al. (2025, 2026b). A one-sentence note that these are recent proposals with limited independent benchmarking (parallel to the CreDE note in Table 3) would match the paper’s care elsewhere and reduce any appearance of uneven scrutiny.
- [Figure 1] Figure 1 is dense; the quality color scale and the closed-form span across scopes are useful but hard to parse in grayscale. A simplified schematic in the main text and a fuller version in the supplement (or clearer legend encoding) would help.
- [Table 1] Table 1 timeline is helpful; a few adjacent items (e.g., conformalized quantile regression, temperature scaling) sit under “Evidential, conformal, and calibration” while core families are year-grouped—consider a single chronological column or clearer subgroup headers for skimming.
- [Section 16] Section 16 is appropriately scoped as adjacent, but the jump from classification ensembles to semantic entropy could cite Malinin & Gales (2021) earlier when first mentioning autoregressive decomposition, to link back to Section 14.
- [Section 3, Eq. (7)] Minor notation: p(w|D) vs p(w| D) spacing is inconsistent in Section 3; EPCE uses H[p,q] for cross-entropy while H also denotes entropy—define cross-entropy explicitly at first use in Eq. (7).
- [Section 1.1, 5.1, title block] Typos/style: “amodeling convention” (Section 1.1) needs a space; “thede factogold standard” (Section 5.1) needs spaces; author email “martin.gillis@.dal.ca” appears to drop a character.
Circularity Check
Survey organization is not circular; only mild non-load-bearing self-citation of the authors’ variance-gated and last-layer work.
specific steps
-
self citation load bearing
[Section 14.7; also Section 7.4 and Table 5/7 (Gillis et al. 2025, 2026a, 2026b)]
"To reduce this cost, Variance-Gated Ensembles (VGEs) (Gillis et al., 2025, 2026b) gate member probabilities by a signal-to-noise factor derived from the ensemble mean and per-class predictive spread, yielding a total/aleatoric/epistemic decomposition at O(MC), a margin-based epistemic score, Variance-Gated Margin Uncertainty (VGMU)..."
Authors cite their own concurrent/recent work when placing VGE/VGMU/VGN (and last-layer committee machines) in the measure/method landscape. This is ordinary survey self-reference, not a load-bearing uniqueness or derivation step: the survey’s main claims (deep ensembles as multi-modal gold standard; method–measure decoupling; diversity–quality link) do not reduce to these citations and remain supported by independent external sources.
full rationale
This is a scoped survey whose central contribution is taxonomic and comparative (method vs. measure separation; generative-mechanism × network-scope placement in Table 2/Figure 1; qualitative L/M/H ratings anchored to external benchmarks in Section 11). It does not claim a first-principles derivation, uniqueness theorem, or fitted-parameter “prediction.” Comparative claims rest on external literature (e.g., Fort et al. 2020; Ovadia et al. 2019; Wood et al. 2023; Zamyatin et al. 2026) and on structural placement of methods, not on closing a definitional loop. Self-citations to Gillis et al. (2025, 2026a, 2026b) appear only when situating variance-gated measures and last-layer committee machines inside the existing landscape; they are not used to force the gold-standard narrative or the taxonomy. Table 3’s heterogeneous/† ratings are a transparency and evidentiary-strength issue, not circularity by construction. No equation reduces a claimed prediction to its own fitted input. Score 1 reflects only routine author self-reference that is not load-bearing.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Predictive uncertainty decomposes into aleatoric and epistemic components via entropy/MI or variance laws, understood as a modeling convention rather than an intrinsic data property (Sections 1.1, 14).
- domain assumption Between-mode functional diversity from full-network explicit members is the primary driver of strong epistemic UQ under distribution shift (Sections 5–6, 12–13).
- ad hoc to paper Qualitative OOD/calibration tiers can be anchored to MSP vs deep-ensemble performance on standard pairs and Ovadia-style shift suites across heterogeneous studies (Section 11, Table 3).
- domain assumption Sample-based measures are freely interchangeable across methods that expose member/posterior samples; closed-form single-pass methods expose only intrinsic signals unless sampled (Figure 1, Section 17).
- standard math Standard information-theoretic and ensemble-diversity identities (predictive entropy, MI, bias–variance–diversity, pairwise KL/JS) hold as used in the cited sources.
read the original abstract
The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures lack principled uncertainty quantification. This survey provides a structured, critical review of methods for Uncertainty Quantification (UQ) in deep learning, scoped to ensemble-based and approximate Bayesian approaches and the measures used to summarize their outputs. Relative to existing UQ surveys, our contribution is depth on efficient ensemble approximations and single-pass methods, and a unified treatment that separates the method producing a predictive distribution from the measure that summarizes its uncertainty. We organize methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer or single-pass approaches. We situate adjacent work on evidential and prior networks, conformal prediction, and post-hoc calibration, together with the decision-time tasks of out-of-distribution detection and selective prediction. For each, we examine theoretical motivation, implementation, empirical performance, and limitations. We then review ensemble diversity theory and uncertainty measures and their decompositions, contrasting the entropy decomposition with pairwise divergence measures, and consolidate evaluation methodology so that our qualitative comparisons share a common basis. We close with a brief treatment of uncertainty in large language models and open research directions, including efficient epistemic measures for classification, last-layer diversity, diversity and calibration under shift, and hybrid architectures.
Figures
Reference graph
Works this paper leans on
-
[7]
Workshop: Bayesian deep learning
URL:https://bayesiandeeplearning.org/2019/ papers/39.pdf. Workshop: Bayesian deep learning. Valdenegro-Toro, M., 2023. Sub-ensembles for fast uncertainty estimation in neural networks, in: International Conference on Computer Vision Workshops (ICCVW), pp. 4119–4127. doi:10.1109/ICCVW60793.2023.00445. Valiuddin, A., van Sloun, R., Viviers, C., de With, P.,...
arXiv 2019
-
[11]
Monod, M., Micheli, A., Bhatt, S., 2025
doi:10.1038/s41598-021-84854-x. Monod, M., Micheli, A., Bhatt, S., 2025. NeuralSurv: Deep survival analysis with Bayesian uncertainty quantification, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–53. URL:https://openreview.net/ pdf?id=c768Z1FwDL. Mukhoti, J., Kirsch, A., van Amersfoort, J., Torr, P.H.S., Gal, Y ., 2023. Deep det...
arXiv 2025
-
[12]
Workshop: Bayesian Deep Learning
URL:https://bayesiandeeplearning.org/2021/ papers/21.pdf. Workshop: Bayesian Deep Learning. Naeini, M.P., Cooper, G.F., Hauskrecht, M., 2015. Obtain- ing well-calibrated probabilities using Bayesian binning, in: AAAI Conference on Artificial Intelligence, pp. 2901–2907. doi:10.1609/aaai.v29i1.9602. Nandy, J., Hsu, W., Lee, M.L., 2021. Towards maximizing t...
arXiv 2021
-
[16]
Workshop: Bayesian deep learning
URL:https://bayesiandeeplearning.org/2021/ papers/28.pdf. Workshop: Bayesian deep learning. van Amersfoort, J., Smith, L., Teh, Y .W., Gal, Y ., 2020. Uncertainty estimation using a single deep deter- ministic neural network, in: International Confer- ence on Machine Learning (ICML), pp. 9690–9700. URL:https://proceedings.mlr.press/v119/ van-amersfoort20a...
-
[18]
ACM Computing Surveys 58, 1–38
A survey on uncertainty quantification of large lan- guage models: Taxonomy, open research challenges, and fu- ture directions. ACM Computing Surveys 58, 1–38. doi:10. 1145/3744238. Smith, L., Gal, Y ., 2018. Understanding measures of uncer- tainty for adversarial example detection. arXiv preprint. doi:10.48550/arXiv.1803.08533. Srivastava, N., Hinton, G....
-
[19]
Function space diversity for uncertainty prediction via repulsive last-layer ensembles, in: International Con- ference on Machine Learning (ICML), pp. 1–16. URL: https://openreview.net/pdf?id=FbMN9HjgHI. Work- shop: Structured Probabilistic Inference & Generative Mod- eling. Tagasovska, N., Lopez-Paz, D., 2019. Single-model uncertain- ties for deep learni...
arXiv 2019
-
[22]
A rigorous link between deep ensembles and (vari- ational) Bayesian methods, in: Conference on Neural In- formation Processing Systems (NeurIPS), arXiv. pp. 1–30. URL:https://openreview.net/pdf?id=eTHawKFT4h& noteId=UBYXOwJMPl. Wilson, A.G., Izmailov, P., 2020. Bayesian deep learning and a probabilistic perspective of generalization, in: Conference on Neu...
arXiv 2020
-
[40]
Rahaman, R., Thiery, A.H., 2021
URL:https://proceedings.mlr.press/v162/ postels22a/postels22a.pdf. Rahaman, R., Thiery, A.H., 2021. Uncertainty quantification and deep ensembles, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–13. URL:https:// openreview.net/pdf?id=wg_kD_nyAF. Ramé, A., Cord, M., 2021. DICE: Diversity in deep en- sembles via conditional redundan...
2021
-
[49]
Xie, J., Ma, Z., Lei, J., Zhang, G., Xue, J.H., Tan, Z.H., Guo, J., 2022
URL:https://www.jmlr.org/papers/volume24/ 23-0041/23-0041.pdf. Xie, J., Ma, Z., Lei, J., Zhang, G., Xue, J.H., Tan, Z.H., Guo, J., 2022. Advanced dropout: A model-free methodology for Bayesian dropout optimization. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 4605–4625. doi:10.1109/TPAMI.2021.3083089. Xiong, M., Hu, Z., Lu, X., Li, Y...
arXiv 2022
-
[1367]
Chau, S.L., Caprio, M., Muandet, K., 2025
URL:https://dl.acm.org/doi/pdf/10.5555/ 3495724.3495839. Chau, S.L., Caprio, M., Muandet, K., 2025. Integral imprecise probability metrics, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–42. URL:https:// openreview.net/pdf?id=KM2XzHq2Rm. Chauhan, V .K., Zhou, J., Lu, P., Molaei, S., Clifton, D.A.,
arXiv 2025
-
[1622]
Boluki, S., Ardywibowo, R., Dadaneh, S.Z., Zhou, M., Qian, X., 2020
URL:https://proceedings.mlr.press/v37/ blundell15.pdf. Boluki, S., Ardywibowo, R., Dadaneh, S.Z., Zhou, M., Qian, X., 2020. Learnable bernoulli dropout for Bayesian deep learning, in: International Conference on Arti- ficial Intelligence and Statistics (AISTATS), pp. 3905–
2020
-
[1914]
Farquhar, S., Kossen, J., Kuhn, L., Gal, Y ., 2024
URL:https://ieeexplore.ieee.org/stamp/ stamp.jsp?arnumber=9956231. Farquhar, S., Kossen, J., Kuhn, L., Gal, Y ., 2024. De- tecting hallucinations in large language models using se- mantic entropy. Nature 630, 625–630. doi:10.1038/ s41586-024-07421-0. Felicioni, N., Maystre, L., Ghiassian, S., Ciosek, K., 2024. On the importance of uncertainty in decision-...
-
[2005]
Active learning for probability estimation using Jensen-Shannon divergence, in: European Conference on Machine Learning, pp. 268–279. doi:10.1007/11564096_ 28. Mena, J., Pujol, O., Vitrià, J., 2021. A survey on uncertainty estimation in deep learning classification systems from a Bayesian perspective. ACM Computing Surveys 54, 1–35. doi:10.1145/3477140. 2...
-
[2015]
Weight uncertainty in neural networks, in: Interna- tional Conference on Machine Learning (ICML), pp. 1613–
-
[2020]
On Last-Layer Algorithms for Classification: Decoupling Representation from Uncertainty Estimation
On last-layer algorithms for classification: Decoupling representation from uncertainty estimation. arXiv preprint. doi:10.48550/arXiv.2001.08049. Brown, G., Wyatt, J., Harris, R., Yao, X., 2005. Diversity cre- ation methods: A survey and categorisation. Information Fu- sion 6, 5–20. doi:10.1016/j.inffus.2004.04.004. Charpentier, B., Zügner, D., Günnemann...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2001.08049 2001
-
[2023]
Post-hoc uncertainty learning using a Dirichlet meta- model, in: Conference on Artificial Intelligence, pp. 9772–
-
[2024]
Artificial Intelligence Review 57
A brief review of hypernetworks in deep learn- ing. Artificial Intelligence Review 57. doi:10.1007/ s10462-024-10862-8. Damianou, A.C., Lawrence, N.D., 2013. Deep gaussian process, in: International Conference on Artificial Intelli- gence and Statistics (AISTATS), pp. 1–9. URL:https: //proceedings.mlr.press/v31/damianou13a.pdf. Daxberger, E., Kristiadi, A...
-
[2061]
Liang, J., Hou, R., Hu, M., Chang, H., Shan, S., Chen, X., 2025
URL:https://proceedings.mlr.press/v70/ li17a/li17a.pdf. Liang, J., Hou, R., Hu, M., Chang, H., Shan, S., Chen, X., 2025. Revisiting logit distributions for reliable out-of- distribution detection, in: Conference on Neural Informa- tion Processing Systems (NeurIPS), pp. 1–33. URL:https: //openreview.net/pdf?id=FLdLPUqnsP. Liang, S., Li, Y ., Srikant, R., 2...
arXiv 2025
-
[2292]
Wood, D., Mu, T., Webb, A.M., Reeve, H.W.J., Luján, M., Brown, G., 2023
URL:https://proceedings.mlr.press/v216/ wimmer23a/wimmer23a.pdf. Wood, D., Mu, T., Webb, A.M., Reeve, H.W.J., Luján, M., Brown, G., 2023. A unified theory of diversity in ensem- ble learning. Journal of Machine Learning Research 24, 1– 26
2023
-
[3467]
Valdenegro-Toro, M., 2019
URL:https://proceedings.mlr.press/v89/ vaicenavicius19a/vaicenavicius19a.pdf. Valdenegro-Toro, M., 2019. Deep sub-ensembles for fast un- certainty estimation in image classification, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–
2019
-
[3640]
Workshop: Mathematics of modern machine learning
URL:https://raw.githubusercontent.com/ mlresearch/v286/main/assets/schweighofer25a/ schweighofer25a.pdf. Workshop: Mathematics of modern machine learning. Schweighofer, K., Aichberger, L., Ielanskyi, M., Klambauer, G., Hochreiter, S., 2023b. Quantification of uncertainty with adversarial models, in: Neural Information Processing Sys- tems (NeurIPS), pp. 1...
-
[3916]
19 Brier, G.W., 1950
URL:https://proceedings.mlr.press/v108/ boluki20a/boluki20a.pdf. 19 Brier, G.W., 1950. Verification of forecasts expressed in terms of probability. Monthly Weather Review 78, 1–3. doi:10. 1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2. Brosse, N., Riquelme, C., Martin, A., Gelly, S., Moulines, E.,
1950
-
[7058]
URL:https://dl.acm.org/doi/pdf/10.5555/ 3327757.3327808. Malinin, A., Gales, M., 2019. Reverse KL-divergence training of prior networks: Improved uncertainty and adversarial robustness, in: Conference on Neural Information Pro- cessing Systems (NeurIPS), pp. 1–12. URL:https:// proceedings.neurips.cc/paper_files/paper/2019/ file/7dd2ae7db7d18ee7c9425e38df1...
arXiv 2019
-
[9781]
Shorinwa, O., Mei, Z., Lidard, J., Ren, A.Z., Majumdar, A.,
doi:10.1609/aaai.v37i8.26167. Shorinwa, O., Mei, Z., Lidard, J., Ren, A.Z., Majumdar, A.,
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.