{"id":"cc06ae2a-98e1-4cf1-95ac-6bcd2ffb3cb7","arxiv_id":"2505.18280","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Applying the R2D2 shrinkage prior to Bayesian neural network weights, with a variational Gibbs inference scheme, improves prediction and out-of-distribution detection in image tasks.","lead":"A new prior for Bayesian neural networks, called R2D2, aims to shrink unimportant weights to zero while protecting strong signals, and a hybrid training algorithm is proposed. Tests on image benchmarks and medical images show improved accuracy and out-of-distribution detection over existing Bayesian network designs.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The variational Gibbs sampler is not derived from a variational objective, so Algorithm 1's output is not shown to approximate the R2D2 posterior or to inherit Theorem 1's contraction rate; compare it against exact HMC on a small BNN.","rationale":"The reader's weakest assumption identifies exactly the most load-bearing gap: the inference algorithm's variational validity is unproven, so the empirical results and Theorem 1 are not connected. I agree with this assessment rather than inventing a weaker or more speculative objection. The paper deserves credit for releasing code, running multiple baselines, and providing a nontrivial extension of the R2D2 prior to BNNs, but those contributions do not close the inferential gap. A hybrid scheme that alternates exact Gibbs conditionals with backpropagation can be a reasonable heuristic, but the paper makes a stronger claim: it says the strategy 'enhances stability and consistency in estimation' and that the posterior is approximated 'more accurately.' Neither property is proven, and the theoretical section proves a contraction rate for the exact posterior, not for the output of Algorithm 1. The concrete HMC comparison would settle whether the concern actually lands: if Algorithm 1 reproduces the exact R2D2 posterior on a tractable problem, the gap is empirical and not damaging; if it diverges, the headline results lose their stated interpretation. The recommended verdict remains CONDITIONAL, so no change from the reader's verdict is needed.","tokens_in":27777,"tokens_out":10088,"duration_ms":83667,"concrete_test":"Use a small R2D2 Bayesian MLP (e.g., 2 hidden layers, 50–100 inputs, n≈500) where NUTS/HMC can deliver reliable exact posterior draws. Run Algorithm 1 on the same model and data. Check (1) whether the posterior weight moments and predictive distributions from Algorithm 1 match HMC within Monte Carlo error; (2) whether the ELBO of Eq. (1) evaluated on Algorithm 1's output reaches the level achieved by standard mean-field VI with the same prior; (3) whether a full Gibbs cycle with w held fixed never decreases the ELBO. If any check fails, the reported results are not attributable to the R2D2 posterior; if all pass, the concern is answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Central claim: the R2D2 prior improves shrinkage and uncertainty, and Theorem 1 gives near-minimax posterior contraction. The bridge between prior and results is Algorithm 1. In Section 4.2, ψ, ω, ξ, and φ are sampled from full conditionals p(ψ|w,σ,...), p(ω|w,φ,ξ,σ), etc., while w and ρ are updated by backpropagating the ELBO in Eq. (1). No variational family for the shrinkage parameters is defined, and the Gibbs draws are conditioned on the current stochastic w rather than on fixed variational expectations. This is not a mean-field VI coordinate ascent: there is no evidence that the alternating updates maximize or even monotonically increase the ELBO. The paper provides no convergence or ELBO guarantee for this hybrid, despite calling it 'variational Gibbs' and claiming to analyze the ELBO. Therefore the empirical posterior estimates in Tables 3, 5–8 are not established as approximations to the R2D2 posterior. Theorem 1, proved in the appendix via Sun et al. [48]'s conditions, concerns the exact posterior under the R2D2 prior and says nothing about the distribution produced by Algorithm 1. Unless the algorithm can be shown to sample or optimize toward that posterior, the empirical wins cannot be attributed to the R2D2 prior, and the theoretical rate is not connected to the method being proposed. This is the paper's load-bearing gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes R2D2-Net, a Bayesian neural network in which the R2D2 shrinkage prior is placed on the weights. The central claims are that this prior has the highest concentration at zero and the heaviest tail among global–local shrinkage priors, so it shrinks irrelevant weights while preserving strong features; that the proposed 'variational Gibbs' algorithm, combining conditional sampling of shrinkage parameters with gradient-based updates of weights, provides accurate posterior approximation; and that the resulting posterior attains a near-minimax contraction rate (Theorem 1). The authors derive analytical KL divergences for part of the variational objective, and they report extensive experiments on simulations, CIFAR-10/100, TinyImageNet, and a medical OOD benchmark, where R2D2-Net often improves classification accuracy and OOD detection over Gaussian, Horseshoe, and spike-and-slab baselines.","tokens_in":28128,"tokens_out":7643,"duration_ms":59436,"significance":"If the claims were established, the paper would be a useful contribution to Bayesian deep learning: it combines a theoretically motivated shrinkage prior with a scalable inference procedure and provides extensive empirical comparisons, with code released at a public repository. The breadth of experiments, including synthetic studies, larger architectures, and a medical imaging OOD task, is a clear strength, as is the attempt to provide an analytical ELBO and posterior concentration analysis. However, the current manuscript does not establish the connection between the proposed inference algorithm and the theoretical posterior concentration result, and the empirical setting does not match the hyperparameter regimes required by the theory. These gaps are load-bearing because the empirical wins are attributed to the R2D2 prior and the variational Gibbs procedure, while Theorem 1 concerns the exact posterior under assumptions not met by the experiments.","major_comments":[{"comment":"The variational Gibbs algorithm is not derived from any variational objective. In Section 4.2, the shrinkage parameters ψ, ω, ξ, and φ are sampled from their full conditional posterior distributions conditioning on the current stochastic weight draws, while w and ρ are updated by backpropagating the ELBO in Eq. (1). No variational family is defined for the shrinkage parameters, and the alternating updates are not shown to increase or even to preserve the ELBO. This is therefore not a mean-field coordinate-ascent procedure, and the output of Algorithm 1 is not established as an approximation of the R2D2 posterior. Consequently, Theorem 1, which is a statement about the exact posterior under the R2D2 prior, is not connected to the distribution produced by the algorithm used in the experiments. The Limitations paragraph in Section 8 acknowledges that the Gibbs sampling may not handle multimodal posteriors, but it does not address this more basic issue. I would need either a derivation of the updates as coordinate ascent on a well-defined variational family with a monotone ELBO, or a comparison against exact HMC on a small BNN showing that Algorithm 1's output tracks the exact posterior.","section":"Section 4.2 / Algorithm 1"},{"comment":"The proof of Theorem 1 is not complete and contains notational and algebraic problems. The proof verifies condition (4) of Theorem 2 by discussing the tail density of the R2D2 prior, but condition (4) is a statement about log(1/π_b), the prior mass or minimum density, not the tail decay; the link between the two is not provided. In the verification of condition (5), the display ends with the inequality 1 − k_n^{a_π} ≤ D_n^{−(1+u)}, which is not derived and appears inconsistent with the preceding algebra, since the left side tends to 1 as k_n → 0. There is also a notation mismatch in Theorem 2, where the terms ar L and ar H are used interchangeably. As written, the appendix does not establish the minimax contraction rate claimed in Theorem 1.","section":"Appendix, Proof of Theorem 1"},{"comment":"The hyperparameters used in the experiments are incompatible with the theoretical assumptions and with the paper's own tail claims. Theorem 1 requires b to satisfy E_n/(L_n log n + log \\bar H)^{1/2} ≲ b ≲ n^α, where E_n is polynomially growing in n, so b must grow with n. The experiments use a fixed b = 0.5 (Supplementary Section .4). Moreover, Table 1 gives the R2D2 tail decay as O(1/|β|^{1+2b}); with b = 0.5 this is O(1/β^2), identical to the Horseshoe tail. Therefore the abstract's and Section 1's claims that the R2D2 prior has the heaviest tail and avoids over-shrinkage compared to the Horseshoe are not realized under the reported default setting. Similarly, with a_π = 0.6 the concentration-at-zero exponent 1−2a_π is negative, so the claimed divergence at zero in Table 1 does not occur. The authors should either choose hyperparameters satisfying the stated rates or qualify the 'highest concentration / heaviest tail' claims to the parameter regimes in which they hold.","section":"Theorem 1 / Supplementary Section .4"},{"comment":"The pseudocode of Algorithm 1 is inconsistent with the text in Section 4.2. Lines 10–14, which sample ω_l, ξ_l, ψ_jl, and T_jl, are inside the inner loop 'for w_jl in w_l do', so the global shrinkage parameters ω_l and ξ_l are resampled once for every weight in the layer, with only the last draw retained. Section 4.2 specifies that ω_l and ξ_l are sampled once per layer. This ambiguity makes the published algorithm impossible to reproduce unambiguously and should be corrected in the pseudocode, with the layer-wise updates moved outside the per-weight loop.","section":"Algorithm 1"}],"minor_comments":[{"comment":"The displayed definition of KL(q∥π) appears to contain typos: it reads KL(q∥π) = E_{q∈Q}[log p(θ|·)] + H[π(θ)], which is not the Kullback–Leibler divergence; it should presumably be E_q[log q(θ)] − E_q[log π(θ)] (or the equivalent entropy form). Please correct this.","section":"Section 2, Eq. (1)"},{"comment":"The notation s_n is used for 'the input dimension of γ*', but s_n was not introduced earlier and the phrase is ambiguous; please define s_n precisely (e.g., the number of nonzero input coordinates or the number of active input dimensions) and relate it to D_n.","section":"Section 5, Condition A.2.2"},{"comment":"In Table 2, the KL divergence for ψ_jl is written in terms of ψ, but the variational posterior is placed on ψ^{-1} (Reciprocal InvGaussian). The appendix derivation later introduces Y = 1/ψ; the table and the derivation should be written in the same variable, and the resulting closed form should be stated explicitly.","section":"Table 2 / Appendix KL derivation"},{"comment":"The text says that in Scenario 3 'shrinkage methods are expected to underperform as they shrink noise features to zeros', but Table 3 shows R2D2-Net performing best in this scenario; please clarify what 'underperform' means here or remove the sentence, which currently contradicts the reported results.","section":"Section 6.1, Scenario 3"},{"comment":"The appendix organization is confusing: Section A and then Sections C and D are used nontransparently, and several places refer only to 'the supplementary materials' without section numbers. Please unify the numbering and make all forward references explicit.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript extends the R2D2 prior of Zhang et al. [60], and one of the present authors is a coauthor of that paper. This is a self-citation, but the BNN application, the variational Gibbs algorithm, and the experiments appear to be independently developed; I do not see this as a disqualifying issue. The main concern is scope and fit: the paper is submitted to a cs.LG venue, and the majority of the theoretical material is imported from statistics literature. The load-bearing gaps identified above can likely be fixed with additional derivations and a small HMC experiment, but the revision is substantial. The repeated references to 'the supplementary materials' suggest the paper may have been formatted for a statistics journal; the authors should adapt the presentation for the current venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The R2D2-Net is a legitimate extension of the R2D2 prior to BNN weights, and the empirical work is more extensive than most BNN prior papers: synthetic regression, CIFAR/TinyImageNet, OOD detection, a medical imaging benchmark, several architectures, and public code. The closed-form KL divergences for the shrinkage parameters are also useful. That part is solid and citable as an application.\n\nThe soft spot is exactly where the stress-test put it: Algorithm 1's Gibbs steps are not derived from a variational objective. The paper updates w and rho by back-propagating the ELBO, but it samples psi, omega, xi, and phi from conditional posteriors given the current w and sigma. That is a hybrid MCMC/VI procedure with no ELBO guarantee and no proof that the output approximates the R2D2 posterior. Theorem 1 is about the exact posterior under the R2D2 prior, not about Algorithm 1, so the contraction rate does not currently attach to the method that produced the tables. Until the authors either prove the alternating updates move toward the ELBO or at least show on a small network that the algorithm's posterior matches HMC, the empirical wins cannot be confidently attributed to the prior. This is a central gap, though not one that sinks the engineering contribution.\n\nTwo smaller issues. First, the \"heaviest tail\" motivation is overstated: Table 1 gives R2D2 tail decay O(beta^{-(1+2b)}), and the default b=0.5 gives the same O(beta^{-2}) as the Horseshoe. Second, the appendix proof of Theorem 1 has notational slips and leans heavily on Sun et al., which is fine, but the presentation needs cleaning.\n\nThe self-citation to Zhang et al. is not a problem here; the prior is correctly cited and the BNN application is new. The discussion also honestly notes that the Gibbs sampler may struggle with multimodal posteriors and that the R2 analogy is loose in DNNs.\n\nBottom line: send it to a serious referee. The empirical and methodological core deserves scrutiny, and the variational gap is addressable. I would ask for a small HMC comparison and a clearer statement of what Algorithm 1 actually approximates.","headline":"A useful BNN prior and inference recipe with strong empirical work, but the variational Gibbs sampler is not shown to be variational, leaving the theory and the experiments disconnected.","tokens_in":28599,"tokens_out":3379,"would_cite":true,"duration_ms":28412,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62G20","68T07","62J07"],"pacs":[],"model":"deepseek-v4-flash","headline":"A new prior for Bayesian neural networks—the R2D2 prior—shrinks irrelevant weights toward zero while preserving the weights that carry strong signals, with a variational Gibbs inference algorithm and a near-minimax contraction guarantee.","keywords":["Bayesian neural network","R2D2 prior","shrinkage prior","variational inference","variational Gibbs","uncertainty estimation","posterior concentration","sparse deep learning"],"falsifier":"Train R2D2-Net on a small synthetic regression with known sparse ground-truth weights and compare its posterior to the Hamiltonian Monte Carlo oracle posterior used by the paper as ground truth: if the variational Gibbs posterior's mean and variance do not track the HMC oracle, or if the ELBO does not monotonically improve during the Gibbs updates, then the algorithm is not performing variational inference and Theorem 1's contraction guarantee does not apply to the implemented procedure.","tokens_in":27613,"feed_emoji":"🎯","tokens_out":5911,"duration_ms":44962,"temperature":0.7,"pith_summary":"The paper proposes R2D2-Net, a Bayesian neural network that places the R2-induced Dirichlet Decomposition (R2D2) prior on every network weight. The central claim is that this prior shrinks weights of irrelevant or noisy features toward zero while leaving weights carrying strong signals intact, which ordinary Gaussian priors cannot do and which lighter-tailed shrinkage priors such as the Horseshoe do too aggressively. To fit the model, the authors develop a variational Gibbs inference algorithm that alternates Gibbs updates of the shrinkage parameters with gradient-based updates of the weights, and they derive closed-form KL divergences for the shrinkage parameters. They also prove that the R2D2 posterior achieves a near-minimax contraction rate under a polynomial-boundedness condition on the true weights. If the claims hold, BNNs using this prior would suffer less variance inflation, support deeper architectures, and give more reliable uncertainty estimates on out-of-distribution inputs.","feed_headline":"BNN prior shrinks noise, spares key features","feed_subtitle":"The R2D2 prior shrinks noisy weights to zero while preserving strong signals, sharpening BNN predictions and uncertainty.","key_machinery":"The carrying object is the R2D2 prior, defined by the scale-mixture representation $\\beta_j \\mid \\psi_j,\\phi_j,\\omega \\sim N(0, \\psi_j \\phi_j \\omega \\sigma^2/2)$ with $\\psi_j \\sim \\mathrm{Exp}(1/2)$, $\\phi \\sim \\mathrm{Dir}(a_\\pi,\\dots,a_\\pi)$, $\\omega \\mid \\xi \\sim \\mathrm{Ga}(a,\\xi)$, $\\xi \\sim \\mathrm{Ga}(b,1)$. Its marginal density decays like $O(|\\beta|^{-(1+2b)})$ in the tails and concentrates at zero like $O(|\\beta|^{-(1-2a_\\pi)})$, which the paper compares favorably against the Horseshoe, Horseshoe+, Dirichlet-Laplace, and generalized double Pareto priors. The inference machinery is a variational Gibbs algorithm that alternately samples $\\psi$, $\\omega$, $\\xi$, and $\\phi$ from their conditional posteriors and updates the reparameterized weight means and variances by back-propagating the ELBO, with closed-form KL divergences for the shrinkage parameters.","core_discovery":"Placing the R2D2 prior on the weights of a Bayesian neural network makes the posterior contract toward sparse solutions while staying concentrated enough on large coefficients to preserve predictive features. The paper states this as: the R2D2 prior has the highest concentration rate at zero among the compared global-local shrinkage priors and the heaviest tail, so irrelevant coefficients are shrunk toward zero and key features are not over-shrunk. The variational Gibbs algorithm treats the conditional posteriors of the shrinkage parameters as variational updates and back-propagates the ELBO through reparameterized weights to learn the per-neuron variances. The theoretical result, Theorem 1, asserts that under sparsity and polynomial-boundedness conditions the posterior contracts in Hellinger distance at a near-minimax rate $\\epsilon_n^2 = O(\\varpi_n^2) + O([r_n L_n \\log n + r_n \\log \\bar H + s_n \\log D_n]/n)$, matching the spike-and-slab rate.","pith_inferences":["The shrinkage-versus-preservation tradeoff suggests R2D2-Net could double as a pruning method: weights whose posterior mass concentrates at zero could be removed after training, yielding compressed networks without a separate sparsification pass.","The variational Gibbs update is not derived from a mean-field variational objective, so a stricter derivation (or a correction) might be needed before the empirical results can be tied to the contraction theorem; one test is to check whether the Gibbs steps decrease the ELBO at every iteration.","If the R2D2 prior's heavy-tail property transfers to attention weights, the same construction could give Bayesian transformers a principled way to prune attention heads, an extension the paper flags as non-trivial.","The per-neuron variance parameter $\\sigma_{jl}$ is learned by back-propagation rather than set by regression MSE; a direct comparison against the original R2D2 setting with layer-shared variance would isolate how much of the gain comes from the prior versus the per-neuron variance."],"forward_implications":["On image classification benchmarks (CIFAR-10, CIFAR-100, TinyImageNet), R2D2-Net reports higher accuracy and AUROC than Gaussian, Horseshoe, and spike-and-slab BNN designs, in some cases matching or exceeding the frequentist baseline.","For out-of-distribution detection, the paper reports that R2D2-Net's entropy-based uncertainty scores outperform the compared Bayesian and non-Bayesian baselines on natural and medical image datasets.","The variational Gibbs updates plus closed-form KL divergences make the ELBO more faithful than the Gaussian-KL approximation used in standard mean-field BNNs, which the paper argues reduces variance inflation.","Theorem 1 gives a near-minimax posterior contraction rate under the R2D2 prior, placing it on par with spike-and-slab priors for sparse deep learning."],"supporting_citations":[{"why":"Supplies the R2D2 prior, its scale-mixture representation, and the Gibbs conditional updates that the paper adapts to neural networks.","marker":"[60]"},{"why":"The Horseshoe-prior BNN that serves as the main shrinkage baseline and the motivating example of over- and under-shrinkage.","marker":"[20]"},{"why":"Establishes the spike-and-slab sparse-DNN posterior contraction framework whose rates Theorem 1 extends to the R2D2 prior.","marker":"[48]"},{"why":"Provides the general-prior convergence-rate theorem that the proof of Theorem 1 verifies for the R2D2 density.","marker":"[47]"},{"why":"Defines the Horseshoe prior whose tail and concentration properties are compared against R2D2.","marker":"[8]"}],"fun_headline_variants":["R2D2 prior shrinks noise, spares key features","Bayesian nets shrink noise with R2D2 prior","R2D2-Net: shrinks noise, spares features","Shrink noise, keep features: R2D2 prior for BNNs","R2D2 prior for BNNs: shrink noise, preserve signals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the hybrid variational Gibbs algorithm, which alternates Gibbs samples of the shrinkage parameters with gradient updates of the weights, actually converges to the target posterior and maximizes the evidence lower bound; the paper does not prove this convergence.","fun_headline_variants_meta":{"raw":{"variants":["R2D2 prior shrinks noise, spares key features","Bayesian nets shrink noise with R2D2 prior","R2D2-Net: shrinks noise, spares features","Shrink noise, keep features: R2D2 prior for BNNs","R2D2 prior for BNNs: shrink noise, preserve signals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000331,"raw_usage":{"total_tokens":1867,"prompt_tokens":992,"completion_tokens":875,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":794}},"tokens_in":608,"tokens_out":875,"duration_ms":6628,"temperature":1.0,"reasoning_tokens":794,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:34:10.327679+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train R2D2-Net on a small synthetic regression with known sparse ground-truth weights and compare its posterior to the Hamiltonian Monte Carlo oracle posterior used by the paper as ground truth: if the variational Gibbs posterior's mean and variance do not track the HMC oracle, or if the ELBO does not monotonically improve during the Gibbs updates, then the algorithm is not performing variational inference and Theorem 1's contraction guarantee does not apply to the implemented procedure.","supporting_citations":[{"cited_title":"Bayesian regression using a prior on the model fit: The r2-d2 shrinkage prior.Journal of the American Statistical Association, pages 1–13, 2020","cited_arxiv_id":null,"evidence_quote":"Supplies the R2D2 prior, its scale-mixture representation, and the Gibbs conditional updates that the paper adapts to neural networks."},{"cited_title":"Model selection in bayesian neural networks via horse- shoe priors.J","cited_arxiv_id":null,"evidence_quote":"The Horseshoe-prior BNN that serves as the main shrinkage baseline and the motivating example of over- and under-shrinkage."},{"cited_title":"Learning sparse deep neural networks with a spike-and-slab prior","cited_arxiv_id":null,"evidence_quote":"Establishes the spike-and-slab sparse-DNN posterior contraction framework whose rates Theorem 1 extends to the R2D2 prior."},{"cited_title":"Consistent sparse deep learning: Theory and computation.Journal of the American Statistical Association, 117(540):1981–1995, 2022","cited_arxiv_id":null,"evidence_quote":"Provides the general-prior convergence-rate theorem that the proof of Theorem 1 verifies for the R2D2 density."},{"cited_title":"Handling sparsity via the horseshoe","cited_arxiv_id":null,"evidence_quote":"Defines the Horseshoe prior whose tail and concentration properties are compared against R2D2."}],"review_version":1}