{"id":"6806ea7d-680b-4d88-9e81-3cb5d1a6b246","arxiv_id":"2502.02032","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"HDBEN jointly estimates mean and log-variance coefficients with Bayesian elastic net priors, claiming posterior concentration, selection consistency, and asymptotic normality.","lead":"This paper proposes HDBEN, a Bayesian regression method that models both the mean and the error variance with elastic net penalties, targeting high-dimensional data where the error spread changes across observations. A generalist might care because it promises simultaneous variable selection in both mean and variance, but the theoretical proofs contain serious gaps.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2 is false as stated: the continuous elastic-net prior assigns zero posterior probability to exact zero coefficients, so the support-recovery event has probability 0 for sparse truth.","rationale":"The reader's weakest assumption identifies the exact-zero issue in Theorem 2, and I agree it is the most load-bearing concern. The paper's headline theoretical contribution is variable selection consistency, and that claim fails at the level of model specification: the continuous elastic-net prior cannot give positive posterior probability to exact zeros. This is an internal inconsistency, not a disagreement with prevailing consensus. Even before considering the further problems with Assumption 4.2 contradicting d > n, Corollary 1 assuming a contraction rate it does not prove, and Theorem 1 omitting the Schwartz test and entropy conditions for growing dimension, the support-recovery theorem is identically false under the stated prior. The model and MCMC scheme may be plausible as empirical tools, and the simulation results may be informative, but empirical performance does not repair a theorem whose conclusion has prior probability zero. The reader's REJECT verdict remains appropriate; my independent read does not change it. I see no need to manufacture an additional objection when this one is decisive.","tokens_in":17363,"tokens_out":3719,"duration_ms":42196,"concrete_test":"Analytic decisive check: specialize to d = 1 with β0 = 0 (and, for concreteness, γ0 = 0) and compute the marginal posterior p(β | y, X) ∝ N(y | Xβ, exp(Xγ0)) × π(β), where π(β) is the Section 3.1 Bayesian elastic-net prior. This posterior has a Lebesgue density, so P(β = 0 | y, X) = 0 for every n, whereas Theorem 2 requires P(Sβ = Sβ0 | y, X) → 1. Running the Section 3.2 Gibbs sampler on this one-coordinate case will likewise produce no exact-zero draws, confirming that the theorem's event cannot occur.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 defines β | τβ, λ2,β ~ N(0, Σβ) with τβ,j | λ1,β ~ Exp(λ1,β^2/2) (Eqs. 3-6), and the same absolutely continuous hierarchy for γ. The likelihood in Eq. 7 is a continuous Gaussian density. Therefore, for any finite n and observed (y, X), the joint posterior of (β, γ) is absolutely continuous with respect to Lebesgue measure. For every j, P(β_j = 0 | y, X) = 0 and P(γ_j = 0 | y, X) = 0. Theorem 2 asserts P(Sβ = Sβ0, Sγ = Sγ0 | y, X) → 1, where Sβ = {j : β_j ≠ 0}. Whenever the true support is sparse, this event requires at least one zero coordinate to be exactly zero, so its posterior probability is identically 0 for every n, not converging to 1. The proof's 'true negatives' step claims exact zeros arise from the L1 penalty, but the hierarchical prior is a scale mixture of normals with no point mass at zero; it shrinks coefficients but cannot produce exact zeros. This is not a minor technical gap: Theorem 2 is the basis of the abstract's variable-selection-consistency claim, and it is internally inconsistent with the model specification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes the Heteroscedastic Double Bayesian Elastic Net (HDBEN), a Bayesian regression model that jointly models the conditional mean and log-variance via hierarchical elastic-net priors. The authors claim posterior concentration, variable selection consistency, and asymptotic normality of the posterior in high-dimensional settings, and they support these claims with simulation studies comparing HDBEN to OLS, Lasso, Elastic Net, Bayesian Lasso, and Bayesian Elastic Net. The manuscript contains four theorems: Theorem 1 asserts posterior concentration, Corollary 1 asserts a near-optimal contraction rate for posterior means, Theorem 2 asserts variable selection consistency, and Theorems 3-4 assert asymptotic normality and a higher-order expansion.","tokens_in":17657,"tokens_out":5930,"duration_ms":60799,"significance":"If the theoretical results were valid, the paper would offer a useful unified Bayesian treatment of heteroscedastic regression with simultaneous shrinkage of mean and variance coefficients. The model specification is natural and the simulation section suggests the method can perform well empirically. However, the theoretical core is not established: Theorem 2 is false as stated for the specified continuous prior, Corollary 1 assumes the rate it claims to prove, and Theorem 1's invocation of Schwartz's theorem omits the required testing and entropy conditions. Because the central claims of posterior concentration, selection consistency, and asymptotic normality are load-bearing for the paper's contribution, the significance of the work as presented is severely diminished.","major_comments":[{"comment":"Theorem 2 asserts P(Sβ = Sβ0, Sγ = Sγ0 | y, X) → 1, where Sβ = {j : βj ≠ 0}. Under the specified hierarchical prior, β and γ have continuous Normal-scale-mixture priors with no point mass at zero, and the likelihood is a continuous Gaussian density. Consequently, for every j, P(βj = 0 | y, X) = 0 and P(γj = 0 | y, X) = 0 for every finite n and every observed (y, X). For a sparse true support, the event {Sβ = Sβ0, Sγ = Sγ0} requires some coefficients to be exactly zero, so its posterior probability is identically zero, not converging to 1. The proof's \"true negatives\" step claims exact zeros arise from the L1 penalty, but the prior shrinks coefficients without producing exact zeros. This is not a technical gap; the theorem is internally inconsistent with the model specification.","section":"§3.1, Eqs. (3)-(6); §4, Theorem 2"},{"comment":"The contraction rate is assumed rather than derived. Theorem 1 only states that there exists a sequence εn → 0 such that the posterior concentrates in an εn-neighborhood; it provides no rate. Corollary 1 then \"sets\" εn = sqrt((sβ+sγ) log d / n) and uses that to obtain the risk bound, and Theorem 2 repeats the same choice. This is circular. Moreover, posterior concentration in probability does not imply the unconditional expectation bound in Eq. (16); without control of the low-probability complement, the expected squared error can be large.","section":"§4, Corollary 1 and Theorem 2"},{"comment":"The proof of Theorem 1 invokes Schwartz's theorem after bounding the Kullback-Leibler divergence locally, but it never verifies the two essential hypotheses: the existence of exponentially consistent tests for the complement of the neighborhood, and prior mass conditions on K-L neighborhoods (the latter are only asserted via Assumption 4.4). In the high-dimensional regime d > n, Assumption 4.2 (full column rank of X as n → ∞) cannot hold, so the setting is internally inconsistent. Thus Theorem 1's conclusion is not established.","section":"§4, Theorem 1 and Assumption 4.2"},{"comment":"The Bernstein-von Mises result is not justified. The theorem assumes a LAN condition and smoothness but does not verify them for the heteroscedastic likelihood; more concretely, the displayed Fisher information matrix has zero off-diagonal blocks between the β and γ sub-vectors, but the log-likelihood's cross-derivatives with respect to β and γ are generally nonzero because the variance exp(Xᵢᵀγ) depends on γ. The limiting covariance in Theorem 3 is therefore incorrect for the stated model.","section":"§4, Theorem 3"},{"comment":"The full conditional for τβ is stated to be inverse Gaussian, but with Σβ = (Dβ^{-1} + λ2,β I)^{-1}, the conditional distribution of β_j given τβ has variance τβ,j/(1 + λ2,β τβ,j), not τβ,j. The joint prior is not the standard Bayesian-lasso scale-mixture representation, so the inverse Gaussian full conditional does not follow. This calls into question the correctness of the MCMC sampler used in the simulations.","section":"§3.2, Eqs. (10)-(11)"}],"minor_comments":[{"comment":"The proof contains a missing citation: \"(see, e.g., ?)\" should be replaced with a proper reference for the K-L divergence between Gaussian densities.","section":"§4, Theorem 1 proof"},{"comment":"Assumption 4.2 states that the design matrix has full column rank as n → ∞, which contradicts the paper's high-dimensional setting where d can exceed n. If the intended assumption is d ≤ n for the asymptotic theory, this should be stated explicitly and reconciled with the simulation settings.","section":"§4, Assumption 4.2"},{"comment":"The theorem heading states \"Assumptions 4.1–4.6 and 4.5,\" which duplicates Assumption 4.5; the intended list should be corrected.","section":"§4, Theorem 4"},{"comment":"The caption refers to \"HDEN β1\" while the method is called HDBEN throughout; the label should be updated.","section":"§5, Figure 2 caption"},{"comment":"Table 1 reports \"Heteroscedasticity γ0,j [0.1, 1.5]\" but the text later says \"γ0 ∈ [10, 100]\"; the values are inconsistent and the notation for the true variance coefficients should be clarified.","section":"§5.1, Table 1"}],"recommendation":"reject","confidential_remarks":"The paper's central theoretical claims are not merely incomplete but internally inconsistent with the model specification (Theorem 2) and rely on circular reasoning (Corollary 1). The empirical section alone, while demonstrating reasonable performance, does not salvage the manuscript's stated contribution. The authors would need to substantially revise the model (e.g., point-mass mixture priors) and rewrite the theoretical analysis to make the claims valid."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the model is a reasonable extension of the Bayesian elastic net to heteroscedastic regression, but the theoretical results are not salvageable as stated. Theorem 2 is false because the continuous hierarchical prior assigns zero posterior probability to exact zero coefficients, so support recovery cannot hold. Corollary 1 simply assumes the contraction rate it is supposed to prove, and Theorem 1 invokes Schwartz without the required test and entropy conditions.\n\nTo be fair, the construction is new relative to the cited literature: Li and Lin (2010) and Park and Casella (2008) only treat the mean. Putting a second elastic-net prior on log-variance is a natural and potentially useful idea, and the Gibbs/MH sampler is standard. The simulations show strong performance, and if code and data were available I would be curious to run them.\n\nBut the proof problems are not cosmetic. Theorem 2's proof explicitly relies on the Laplace penalty producing exact zeros, which it cannot in a prior specified as a scale mixture of normals with no point mass at zero. The KL bound in Theorem 1 also contains a linear term in ||γ−γ0|| that cannot be absorbed into a quadratic bound for small neighborhoods. And Assumption 4.2, full column rank, contradicts the d > n regime the paper claims to address. On the empirical side, there is no code, no data, and no hyperparameter values; the reproducibility note only mentions seeded RNGs. There is even a \"see, e.g., ?\" placeholder in the proof of Theorem 1, which suggests the manuscript is not ready.\n\nWho this is for: readers working on heteroscedastic Bayesian regularized regression might take the model idea and redo the theory, but they should not cite these theorems. My recommendation: do not send to a serious referee yet. This should be desk rejected with an invitation to resubmit after a complete rewrite of the theory section and release of code and data. The underlying model is worth another look once the claims are brought in line with what the prior actually does.","headline":"A promising model idea undermined by unsound theory: the paper's central variable-selection theorem contradicts its own continuous prior, and the contraction results are assumed rather than proved.","tokens_in":18187,"tokens_out":4244,"would_cite":false,"duration_ms":41610,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62J07","62J05","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"HDBEN jointly models the mean and log-variance with Bayesian elastic-net priors, claiming posterior concentration, exact support recovery, and asymptotic normality for both coefficient vectors under mild conditions.","keywords":["Heteroscedasticity","Bayesian Elastic Net","Posterior Concentration","Variable Selection","High-Dimensional Regression","Log-variance model","Scale mixture of normals","MCMC"],"falsifier":"Compute the posterior probability $P(\\beta_j = 0 \\mid y, X)$ in the stated model for any finite $n$ and $d$; it is zero for every $j$ because the posterior density of each coefficient is continuous, directly contradicting the claim that $P(S_\\beta = S_{\\beta_0}, S_\\gamma = S_{\\gamma_0} \\mid y, X)$ tends to one. A simulation counterpart is to run the described Gibbs sampler on data with sparse true supports and record whether any MCMC draw of $\\beta$ or $\\gamma$ ever equals zero exactly.","tokens_in":17085,"feed_emoji":"📊","tokens_out":7597,"duration_ms":69810,"temperature":0.7,"pith_summary":"This paper proposes the Heteroscedastic Double Bayesian Elastic Net (HDBEN), a Bayesian regression model that fits the conditional mean and the conditional log-variance together, with an elastic-net-style prior on both coefficient vectors. The central claim is that this dual regularization simultaneously selects predictors for the mean and for the variance even when the number of predictors exceeds the sample size. If the claim is right, a single fitted posterior yields both point estimates and uncertainty statements that adapt to predictor-dependent error variability. The authors state asymptotic guarantees: posterior concentration, exact support recovery, and asymptotic normality of active coefficients, supported by simulation comparisons against OLS, Lasso, elastic net, Bayesian Lasso, and Bayesian elastic net.","feed_headline":"One Bayesian model recovers sparse mean and variance coefficients","feed_subtitle":"Joint elastic-net priors on mean and log-variance promise near-optimal concentration and selection in high dimensions.","key_machinery":"The carrying mechanism is the hierarchical Bayesian elastic-net prior placed on both coefficient vectors: $\\beta \\mid \\tau_\\beta$ is normal with covariance $(D_\\beta^{-1} + \\lambda_{2,\\beta} I)^{-1}$, and each $\\tau_{\\beta,j}$ follows an exponential distribution with rate $\\lambda_{1,\\beta}^2/2$, with the identical construction for $\\gamma$. This scale-mixture-of-normals representation is the Bayesian analogue of the combined $\\ell_1+\\ell_2$ penalty, giving sparsity through the exponential mixing and stable grouped shrinkage through the ridge-type precision. In HDBEN it is applied twice, once for the mean and once for the log-variance, so that variance coefficients can be selected and shrunk in the same sense as mean coefficients. The posterior sampler exploits conditional conjugacy: $\\beta$ has a normal full conditional, $\\tau_\\beta$ and $\\tau_\\gamma$ have inverse-Gaussian full conditionals, the shrinkage hyperparameters have Gamma full conditionals, and $\\gamma$ is updated with a Metropolis-Hastings step.","core_discovery":"The paper's central claim is that applying the Bayesian elastic-net hierarchy to the log-variance coefficients $\\gamma$, alongside the usual mean coefficients $\\beta$, turns the heteroscedastic Gaussian model $y_i \\sim N(X_i^T\\beta, \\exp(X_i^T\\gamma))$ into a jointly regularized Bayesian procedure. HDBEN's posterior is claimed to concentrate around the true parameters at a rate of order $\\sqrt{(s_\\beta+s_\\gamma)\\log d / n}$, to recover the exact sparsity patterns of both $\\beta$ and $\\gamma$ with probability tending to one, and to be asymptotically normal on the active coefficients under mild regularity conditions. The same scale-mixture prior that gives the Bayesian elastic net its $\\ell_1+\\ell_2$ behavior is placed on both parameter blocks, so sparsity and grouping are induced in the variance model as well as the mean model. Simulations with $d$ up to 1000 report lower estimation error for $\\beta$ than the compared baselines, particularly where heteroscedasticity is strong.","pith_inferences":["Extension: because the normal-exponential hierarchy gives every coefficient a continuous posterior density, the exact-support consistency claim in Theorem 2 would need a spike-and-slab prior or a post-hoc thresholding rule to be literally true.","Extension: the random-walk Metropolis update for the $d$-dimensional vector $\\gamma$ is likely to mix more slowly as $d$ grows; a testable remedy is a reparameterized or adaptive sampler, and effective sample sizes should be compared across dimensions.","Extension: the same dual elastic-net construction transfers to heteroscedastic GLMs where both location and dispersion depend on predictors, since the variance-coefficient prior does not rely on Gaussian responses.","Extension: a finite-sample check of support recovery against thresholded HDBEN posteriors and a spike-and-slab competitor would show how much of the claimed selection consistency is visible in practice."],"forward_implications":["Users obtain one posterior that estimates both mean effects and variance effects, so prediction intervals can widen automatically where the predictors indicate higher volatility.","Under the stated theory, the squared-error risk of the posterior means decays as $O((s_\\beta+s_\\gamma)\\log d / n)$, the usual near-optimal high-dimensional rate.","If exact support recovery holds, the fitted model tells the practitioner which covariates shift the response and which covariates change its variability, with no separate selection step.","Asymptotic normality of the active coefficients would justify reporting posterior-based confidence intervals for the selected mean and variance parameters."],"supporting_citations":[{"why":"Defines the frequentist elastic net whose combined $\\ell_1$ and $\\ell_2$ penalties the Bayesian hierarchy mirrors.","marker":"Zou and Hastie, 2005"},{"why":"Supplies the hierarchical Bayesian elastic-net prior that HDBEN extends to the variance coefficients.","marker":"Li and Lin, 2010"},{"why":"Provides the normal-exponential scale-mixture representation used for Bayesian lasso-style shrinkage in the sampler.","marker":"Park and Casella, 2008"},{"why":"Gives the posterior-consistency theorem invoked to prove concentration around the true parameters.","marker":"Schwartz, 1965"},{"why":"Defines the Lasso baseline whose selection behaviour the paper's high-dimensional claims build on.","marker":"Tibshirani, 1996"}],"fun_headline_variants":["Joint Bayesian elastic net for sparse mean and variance","Heteroscedastic data? HDBEN regularizes both layers","Posterior concentration for mean and variance coefficients","Double elastic net: sparsity in regression and heteroscedasticity","Sparse Bayesian joint model for mean and log-variance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a continuous Bayesian elastic-net prior can produce exact zeros in the posterior, yet under the paper's normal-exponential hierarchy every coefficient has a continuous posterior density, so the event of exact support recovery has posterior probability zero.","fun_headline_variants_meta":{"raw":{"variants":["Joint Bayesian elastic net for sparse mean and variance","Heteroscedastic data? HDBEN regularizes both layers","Posterior concentration for mean and variance coefficients","Double elastic net: sparsity in regression and heteroscedasticity","Sparse Bayesian joint model for mean and log-variance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000523,"raw_usage":{"total_tokens":2515,"prompt_tokens":915,"completion_tokens":1600,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":1518}},"tokens_in":531,"tokens_out":1600,"duration_ms":12696,"temperature":1.0,"reasoning_tokens":1518,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T13:36:44.151959+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the posterior probability $P(\\beta_j = 0 \\mid y, X)$ in the stated model for any finite $n$ and $d$; it is zero for every $j$ because the posterior density of each coefficient is continuous, directly contradicting the claim that $P(S_\\beta = S_{\\beta_0}, S_\\gamma = S_{\\gamma_0} \\mid y, X)$ tends to one. A simulation counterpart is to run the described Gibbs sampler on data with sparse true supports and record whether any MCMC draw of $\\beta$ or $\\gamma$ ever equals zero exactly.","supporting_citations":[{"cited_title":"and Hastie, T","cited_arxiv_id":null,"evidence_quote":"Defines the frequentist elastic net whose combined $\\ell_1$ and $\\ell_2$ penalties the Bayesian hierarchy mirrors."},{"cited_title":"and Casella, G","cited_arxiv_id":null,"evidence_quote":"Provides the normal-exponential scale-mixture representation used for Bayesian lasso-style shrinkage in the sampler."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the posterior-consistency theorem invoked to prove concentration around the true parameters."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Lasso baseline whose selection behaviour the paper's high-dimensional claims build on."}],"review_version":1}