{"id":"20175126-7ada-4677-8a9f-ffbd238cdd4e","arxiv_id":"2607.28248","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured survey separates UQ methods from uncertainty measures, rates ensemble and single-pass families on cost and quality, and flags open gaps in efficient epistemic scoring and last-layer diversity.","lead":"This survey organizes deep-learning uncertainty methods by how they produce a predictive ensemble and separately by how uncertainty is measured from that ensemble. It is useful as a practitioner map of cost–quality trade-offs among ensembles, Bayesian approximations, and single-pass detectors.","discovery_kind":"review","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The paper is an explicit survey (discovery_kind=review), not a new empirical result. Its load-bearing contribution is the method/measure separation and the two-axis taxonomy that organizes why multi-modal full-network exploration outperforms efficient approximations on shift/OOD while allowing distance-aware single-pass methods to punch above their mechanism class on OOD. The reader correctly flags Table 3's qualitative ratings as the softest evidentiary link; the paper itself states they are relative placements from heterogeneous benchmarks and marks structure-inferred cells with daggers. That limitation is normal for surveys and is not required to be a unified re-benchmark for the organizational claim to hold. Diversity theory and the cited head-to-head studies (Ovadia, Fort, Gustafsson, Zamyatin, Liu, Mukhoti) independently support the gold-standard narrative and the sacrifice-of-diversity pattern. No equation-level error, scope overclaim, or undisclosed contradiction appears. Verdict remains ACCEPT; no adjustment warranted.","tokens_in":38913,"tokens_out":589,"duration_ms":11460,"concrete_test":"Spot-check three Table 3 cells against their cited Evidence sources (deep ensemble H/H from Ovadia et al. 2019; BatchEnsemble L/L from Zamyatin et al. 2026; SNGP M–H from Liu et al. 2022) and confirm the paper's own caveats in Section 11–12 match the source limitations; if those cells are faithful summaries and the gold-standard narrative is hedged as written, the central claim stands without adjustment.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption (heterogeneous L/M/H ratings in Table 3) is real but already disclosed by the paper and does not undermine the central claim. The strongest claim is organizational and comparative: method vs. measure separation plus the generative-mechanism × network-scope taxonomy (Table 2, Figure 1) explain why full-network multi-modal methods set the quality standard while shared-weight/single-basin approximations lose functional diversity, with distance-aware single-pass methods as a partial exception on OOD. That claim rests primarily on the loss-landscape and diversity literature (Fort et al. 2020; Wood et al. 2023; Zamyatin et al. 2026) and on the structural placement of methods, not on treating the L/M/H cells as precise cardinal rankings. Section 11 and Table 3 note (a)/† explicitly frame the ratings as relative, evidence-anchored placements from heterogeneous studies. For a scoped survey this is standard and does not create an internal inconsistency or a hidden technical error.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This survey reviews uncertainty quantification for deep learning, scoped to ensemble-based and approximate Bayesian methods and to the measures that summarize their predictive ensembles. Its organizing contribution is a clean separation between the method that produces a predictive distribution and the measure that reads uncertainty from it, together with a two-axis taxonomy (generative mechanism × network scope; Table 2, Figure 1). Methods are grouped into BNNs, MC Dropout, deep ensembles, efficient ensemble approximations, and last-layer/single-pass approaches, with adjacent treatment of evidential/prior networks, conformal prediction, post-hoc calibration, OOD detection, selective prediction, and a brief LLM section. The authors consolidate evaluation practice (Section 11), diversity theory (Section 13), and entropy versus pairwise-divergence measures (Section 14), and close with open directions on last-layer diversity, efficient epistemic measures, shift, hybrids, and generative UQ.","tokens_in":39147,"tokens_out":1363,"duration_ms":31746,"significance":"If the organizational claims hold—as the loss-landscape and diversity literature largely supports—the paper offers a usable map of a crowded field that existing surveys treat more broadly or more architecture-first. Depth on efficient ensembles (BatchEnsemble, SWAG/MultiSWAG, CreDE, masksembles) and single-pass/last-layer methods (SNGP, DUQ/DDU/DUE, LLE, VBLL), plus an explicit method–measure decoupling and a consolidated evaluation basis for qualitative comparisons, are genuine contributions relative to Gawlikowski et al., Abdar et al., and He et al. The structural reading that full-network multi-modal exploration sets the quality standard while shared-weight/single-basin approximations lose functional diversity, with distance-aware single-pass methods as a partial OOD exception, is a clear and citable synthesis. Open problems (efficient non-entropy epistemic measures, last-layer diversity limits, hybrids, LLM semantic UQ) are well posed for follow-on work.","major_comments":[{"comment":"Section 11–12 and Table 3 note (a)/†: the L/M/H OOD and calibration ratings are the main comparative device supporting family rankings, yet they pool heterogeneous architectures, datasets, and shift types, with several † ratings inferred from structure rather than direct benchmarks. The manuscript already discloses this and frames the cells as relative placements, which is appropriate for a survey; still, the gold-standard narrative and several M/H contrasts would be more robust if Table 3 (or a short appendix) marked which cells rest on OpenOOD/Uncertainty Baselines-style head-to-heads versus originating-paper or structural inference, so readers can weight the rankings accordingly.","section":"Section 11–12, Table 3"},{"comment":"Section 14.5–14.7 and Tables 6–8: the structural case that EPJS removes EPKL’s unboundedness failure mode while avoiding MI’s mixture underestimation is carefully stated as non-empirical, and validation on standard OOD/shift suites is correctly listed as open (Section 17.2). Because the survey’s measure contribution partly rests on elevating pairwise and variance-gated alternatives over the default MI decomposition, a brief explicit caveat in Section 14.5 (and in Table 7’s EPJS/VGMU rows) that no claim of superior detection AUROC is being made would prevent over-reading by practitioners.","section":"Section 14.5–14.7, Tables 6–8"}],"minor_comments":[{"comment":"Section 14.7 and Table 5/7: variance-gated ensembles, VGMU, and VGN are introduced largely via Gillis et al. (2025, 2026b). A one-sentence note that these are recent proposals with limited independent benchmarking (parallel to the CreDE note in Table 3) would match the paper’s care elsewhere and reduce any appearance of uneven scrutiny.","section":"Section 14.7"},{"comment":"Figure 1 is dense; the quality color scale and the closed-form span across scopes are useful but hard to parse in grayscale. A simplified schematic in the main text and a fuller version in the supplement (or clearer legend encoding) would help.","section":"Figure 1"},{"comment":"Table 1 timeline is helpful; a few adjacent items (e.g., conformalized quantile regression, temperature scaling) sit under “Evidential, conformal, and calibration” while core families are year-grouped—consider a single chronological column or clearer subgroup headers for skimming.","section":"Table 1"},{"comment":"Section 16 is appropriately scoped as adjacent, but the jump from classification ensembles to semantic entropy could cite Malinin & Gales (2021) earlier when first mentioning autoregressive decomposition, to link back to Section 14.","section":"Section 16"},{"comment":"Minor notation: p(w|D) vs p(w| D) spacing is inconsistent in Section 3; EPCE uses H[p,q] for cross-entropy while H also denotes entropy—define cross-entropy explicitly at first use in Eq. (7).","section":"Section 3, Eq. (7)"},{"comment":"Typos/style: “amodeling convention” (Section 1.1) needs a space; “thede factogold standard” (Section 5.1) needs spaces; author email “martin.gillis@.dal.ca” appears to drop a character.","section":"Section 1.1, 5.1, title block"}],"recommendation":"minor_revision","confidential_remarks":"Fit for a methods/survey venue in stat.ML or ML is good. The self-citation cluster on variance-gated measures and last-layer committee machines is noticeable but secondary to the taxonomy and is handled more carefully than many surveys handle authors’ own lines; no need to block on it if the minor caveat is added. I agree with the external stress-test that the heterogeneous Table 3 ratings are a real limitation but already disclosed and not load-bearing against the organizational claim."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that this is a careful organizational survey, not a new method paper. What it actually adds is a clean separation between the thing that produces a predictive ensemble and the measure that summarizes it, plus a two-axis taxonomy (generative mechanism × network scope) that makes the cost–quality pattern legible: full-network multi-modal methods set the quality bar; shared-weight and single-basin approximations lose functional diversity; distance-aware single-pass methods can still do strong OOD without between-mode diversity.\n\nThat framing is useful. Depth on efficient ensembles (BatchEnsemble, SWAG/MultiSWAG, CreDE, masksembles) and last-layer/single-pass work (SNGP, DUQ/DDU/DUE, LLE, VBLL) is better than the broad surveys they position against. The diversity section (Wood et al., Fort et al.) and the measure discussion (MI vs EPKL/EPJS, Wimmer critique, variance-gated stuff) are fair and current. Limitations are stated in proportion. Citation pattern looks solid; self-cites on VGE/VGMU and last-layer committees sit inside the measure landscape rather than propping the main claim.\n\nSoft spot, already disclosed: Table 3’s L/M/H OOD/calibration cells mix heterogeneous studies and some structure-inferred † ratings. Section 11 and the table note say so. That weakens any reading of the cells as precise cardinal rankings, but the central comparative story does not rest on treating them that way—it rests on loss-landscape and diversity evidence plus structural placement. For a survey this is normal, not load-bearing failure. No unified re-benchmark, no new theorem; that is what the paper is.\n\nWho it is for: people who pick ensembles vs single-pass methods, or who pair methods with measures and OOD/selective protocols. Worth a serious referee. I would bring it to reading group and cite the taxonomy and open problems. Send to peer review.","headline":"Solid scoped survey: method–measure split and mechanism×scope taxonomy are the real value; L/M/H tables are disclosed qualitative placements, not a hidden flaw.","tokens_in":39740,"tokens_out":501,"would_cite":true,"duration_ms":10361,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Uncertainty quality tracks multi-modal diversity: methods that explore many loss basins set the standard, while shared-weight shortcuts and single-pass tricks trade that diversity away.","keywords":["uncertainty quantification","epistemic uncertainty","ensemble diversity","calibration","out-of-distribution detection","deep ensembles","single-pass uncertainty","large language models"],"falsifier":"On a fixed architecture and the standard OpenOOD / shift suites, show that a cheap shared-weight or last-layer method matches a deep ensemble on both calibration under corruption severity and OOD AUROC once compute is matched; or show that last-layer diversity alone recovers most of the ensemble’s epistemic gap.","tokens_in":39779,"feed_emoji":"🎯","tokens_out":1049,"duration_ms":21298,"temperature":0.7,"pith_summary":"Deep networks used in safety-critical settings need honest confidence scores, not just point predictions. This survey argues that the field is clearer once you split two stages that papers usually mix: the method that produces a set of predictions, and the measure that turns that set into total, aleatoric, or epistemic uncertainty. Along a simple grid of how the ensemble is made (explicit members, posterior samples, or a closed-form single pass) and how much of the network is free (full net, last layer, or one deterministic model), the pattern is consistent. Full-network methods that land in different loss basins—deep ensembles and their multi-modal cousins—remain the quality benchmark for calibration and out-of-distribution detection. Efficient approximations that share weights, stay in one basin, or read uncertainty from a single forward pass cut cost but usually cut functional diversity, and with it shift robustness; distance-aware single-pass designs can still score well on OOD without true between-mode diversity. The same split exposes open gaps: better last-layer diversity, cheap epistemic scores for classification that do not rely on fragile entropy identities, hybrids that add a little multi-modal diversity to strong single-pass backbones, and principled uncertainty for free-form language generation.","feed_headline":"Multi-basin ensembles still set the UQ standard","feed_subtitle":"A method–measure split shows shared-weight shortcuts buy speed by sacrificing diversity under shift","key_machinery":"The method–measure separation plus the two-axis taxonomy (how the predictive ensemble is generated: explicit members, posterior sampling, or closed-form predictive; and network scope: full network, shared backbone/last layer, or single deterministic net). That grid, not architecture brand names, is what organizes the quality–cost trade-off and the comparison tables.","core_discovery":"Reading uncertainty quantification through a method–measure split and a generative-mechanism × network-scope taxonomy shows that between-mode, full-network exploration is what delivers top-tier uncertainty, while efficient shared-weight and single-basin methods systematically lose functional diversity and therefore weaker behavior under shift—yet carefully designed distance-aware single-pass models can still reach strong OOD detection without that multi-modal component.","pith_inferences":["If diversity is the real scarce resource, evaluation suites should report effective ensemble size or representation disagreement under shift, not only ECE and AUROC.","Safety cases that rely on single-pass deterministic uncertainty may be fine for far-OOD flagging yet still miss the multi-solution ignorance that full ensembles capture.","The method–measure split suggests a modular software stack: one backend that emits member predictives, and swappable measure plugins—including conformal wrappers—without retraining.","Epistemic collapse at large scale implies that “more parameters” alone will not buy trustworthy uncertainty; explicit multi-modal or repulsive mechanisms may stay necessary."],"forward_implications":["Practitioners should treat deep ensembles (and multi-modal extensions) as the quality reference and judge cheaper methods by how much functional diversity they keep under shift.","Any sample-based ensemble can be paired with any sample-based uncertainty measure; closed-form single-pass methods only expose their built-in signal unless you sample them.","Hybrid designs—small ensembles of distance-aware backbones, or one distance-aware trunk with several heads—become the natural next architecture class to test.","Bounded pairwise scores and linear-time gated epistemic measures need the same standardized OOD/shift bake-off that methods already get.","Generative language uncertainty must redefine diversity and the aleatoric/epistemic split over meanings, not fixed labels."],"fun_headline_variants":["Multi-basin ensembles still lead UQ under shift","Method–measure split: full-network diversity wins","Shared-weight UQ shortcuts lose diversity under shift","Between-mode exploration sets the UQ standard","Single-pass UQ detects OOD without multi-basin depth"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"That qualitative high/medium/low ratings stitched from different papers, architectures, datasets, and shifts—including some inferred only from method structure—are solid enough to rank whole method families head to head.","fun_headline_variants_meta":{"raw":{"variants":["Multi-basin ensembles still lead UQ under shift","Method–measure split: full-network diversity wins","Shared-weight UQ shortcuts lose diversity under shift","Between-mode exploration sets the UQ standard","Single-pass UQ detects OOD without multi-basin depth"]},"model":"grok-4.5","effort":"low","cost_usd":0.003452,"raw_usage":{"total_tokens":1139,"prompt_tokens":791,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":34524000,"prompt_tokens_details":{"text_tokens":791,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":287,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":791,"tokens_out":61,"duration_ms":5328,"temperature":1.0,"reasoning_tokens":287,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T13:22:16.578067+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a fixed architecture and the standard OpenOOD / shift suites, show that a cheap shared-weight or last-layer method matches a deep ensemble on both calibration under corruption severity and OOD AUROC once compute is matched; or show that last-layer diversity alone recovers most of the ensemble’s epistemic gap.","supporting_citations":[],"review_version":1}