REVIEW 3 major objections 5 minor 43 references
Robust Bayesian high-dimensional variable selection and inference with the horseshoe family of priors
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Using a Laplace working likelihood and Gibbs samplers, the paper develops robust Bayesian regression with horseshoe, horseshoe+, and regularized horseshoe priors and reports that the one-group priors can yield calibrated marginal credible…
desk verdict Useful new Gibbs samplers and a serious simulation study, but the abstract overclaims family-level valid inference: regularized horseshoe undercovers badly (0.649) in the paper's own tables. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Laplace likelihood's representation as a normal-exponential scale mixture, combined with the auxiliary-mixture representation of half-Cauchy priors as inverse-gamma scale mixtures. This turns the full posterior into a tractable Gibbs sampler where each coefficient's full conditional is normal, with a shrinkage factor $\kappa_j = (1 + \lambda^2 s_j^2 \sum_i x_{ij}^2 / (\tau^{-1} \xi^2 \tilde v_i))^{-1}$ governing how much the posterior mean is pulled toward zero. Propositions 1.1 through 1.3 express the posterior mean as $(1 - \kappa_j)$ times a least-squares-like update and give the density of $\kappa_j$ for horseshoe and horseshoe+, which is how the paper connects one-group shrinkage to feature selection.
What would settle it
A decisive check is to recompute the empirical coverage of the 95% credible interval for the largest nonzero coefficient under t(2) errors with n = 200 and p = 1000, but with covariates drawn from an independent or banded correlation matrix instead of AR(1); the paper's claim predicts coverage near 0.95, and a coverage below 0.90 for RBHS would falsify the general claim.
Extended reading notes
Core claim
The paper's central claim is that exact sparsity is not necessary for valid Bayesian inference in robust high-dimensional sparse regression. Under a Laplace working likelihood for the errors, horseshoe-family priors shrink noise coefficients so aggressively that effective sparsity emerges, and the resulting marginal credible intervals for the nonzero coefficients achieve empirical coverage near the nominal 95% level. The paper supports this with simulations at (n,p) = (100,500) and (200,1000), reporting that RBHS and RBHS+ maintain coverage roughly from 0.93 to 0.97 under normal and t(2) errors, while RBRHS falls to values such as 0.649 for the largest coefficient in the (200,1000) setting.
Load-bearing premise
The load-bearing premise is that a Laplace working likelihood paired with a horseshoe prior produces posterior credible intervals that hit their nominal coverage when there are more predictors than observations, even though published work cited by the paper shows a variance correction is needed for that likelihood in low dimensions and no correction is available in high dimensions.
Editorial extensions
If this is right
- RBHS and RBHS+ provide a computationally cheap robust alternative to spike-and-slab models: 10,000 Gibbs iterations complete in seconds, with calibration maintained when the number of predictors far exceeds the sample size.
- Under heavy-tailed t(2) errors, the robust horseshoe methods keep coverage near nominal while non-robust counterparts drop to roughly 0.7 to 0.8, so the robustness of the likelihood is what restores interval validity.
- The proposed samplers are orders of magnitude faster than slice-sampling horseshoe quantile regression, making large-scale eQTL-style analyses practical.
- In the rat eye eQTL case study, robust horseshoe models select fewer genes and achieve lower prediction error than non-robust versions.
Reading between the lines
- If the calibration finding extends beyond AR(1) designs, robust high-dimensional inference could drop two-group spike-and-slab priors entirely, removing the need to tune a mixture indicator and lowering computational cost.
- The paper's own RBRHS coverage failures suggest that adding a Gaussian ridge-like component to a horseshoe prior, while helpful for correlated predictors, can break interval calibration; testing whether the b prior or the extra shrinkage layer is responsible would be a direct follow-up.
- A formal theory showing when continuous global-local priors yield calibrated credible intervals under a working likelihood, analogous to the spike-and-slab results, is the natural next step and would turn the simulation finding into a theorem.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops robust Bayesian regression models using the horseshoe, horseshoe+, and regularized horseshoe priors under a Laplace working likelihood, with explicit Gibbs samplers provided in Appendix C. It evaluates variable selection, estimation, credible interval coverage, computational speed, and a real eQTL application, and claims that the one-group horseshoe family can yield valid Bayesian credible intervals in high-dimensional robust regression even without exact sparsity.
Significance. If the family-level inference claim were established, this would be a useful counterpoint to the view that exact sparsity is necessary for valid high-dimensional Bayesian inference, and the explicit Gibbs samplers would be a practical contribution over slice sampling. The simulation infrastructure is solid: the coverage study uses 1,000 replicates, a DGP matching Fan et al. (2024) for cross-comparison, and external comparators such as HSBQR and Bayesreg. The paper is also honest about the absence of a variance-correction theorem for p>n. However, the paper's own results do not support the family-level inference claim: RBRHS coverage drops to 0.649 in Table 15, so the evidence supports RBHS and RBHS+ in a narrow AR(1) design rather than the whole horseshoe family.
major comments (3)
- [Abstract; §3 (Table 4); Appendix A.2 (Table 15)] The abstract and Section 5 claim that 'one-group horseshoe priors' yield valid credible intervals even without exact sparsity, but the coverage tables include RBRHS, which is part of that family. In Table 4, RBRHS coverage for β3 is 0.853 under N(0,1) and 0.805 under t(2); in Table 15 these are 0.781 and 0.649, while RBHS and RBHS+ remain at 0.923–0.971. The family-level statement is therefore contradicted by the paper's own evidence; please restrict the validity claim to RBHS and RBHS+, or provide a method-specific calibration argument for RBRHS, and revise the abstract, contribution list, and Discussion accordingly.
- [§2.2 (Yang et al. (2016) remark; §3 high-dimensional inference)] The paper acknowledges that Yang et al. (2016) requires a posterior variance correction for the Laplace working likelihood and that this correction is unavailable when p>n, and it then relies on the horseshoe prior to repair calibration. No theorem, finite-sample bound, or asymptotic argument is supplied to replace the correction. Given that RBRHS fails the coverage check, the inference claim rests on empirical calibration in one AR(1) design with three strong signals. Please either provide theoretical conditions under which the resulting intervals are calibrated, or explicitly reframe the inference conclusions as empirical and scope them to the settings simulated.
- [Appendix C.1.3 (and C.2.3)] The inverse-Gaussian update for \tilde v_i appears inconsistent with the completed-square derivation. For C.1.3, the displayed conditional has exponent −τ\tilde v_i − τ r_i²/(2ξ² \tilde v_i); under the standard inverse-Gaussian parameterization, the stated parameters μ=√(2ξ²/r_i²), λ=2τ give exponent −τ r_i² \tilde v_i/(2ξ²) − τ/\tilde v_i, which has the roles of \tilde v_i and 1/\tilde v_i swapped. Please check whether the displayed sampler is a reciprocal transformation or a typo; as written, the sampler does not target the stated full conditional. The same issue appears in C.2.3.
minor comments (5)
- [§3 and Table 4] The notation β1 and β2 is used both for scalar coefficients and for the nonzero/zero blocks, and Table 4's 'Coverage of β2' row actually refers to zero coefficients, not coefficient β2=1.5; please use unambiguous notation.
- [Tables 4 and 17] Tables 4 and 17 report the same (n,p)=(100,500), AR(1) setting but give different coverage for the proposed methods (e.g., RBHS β2: 0.940 vs 0.924); please clarify whether these are independent simulation runs or correct the inconsistency.
- [Throughout] There are repeated typos, including 'repametrization', 'an Bayesian', 'multicolinearity', 'proportions' for 'propositions', and 'dintinct meachnisms'; a careful proofread is needed.
- [§3, after Table 9] The text says false positives are 'very low, often close to zero', but for the log-normal non-i.i.d. setting all six methods report about one false positive on average; please adjust the summary sentence to match the table.
- [Main text and Appendix numbering] Appendix A.2 and Table 15 are referenced as 'Figures 5 and 6 and Table 15 in the Appendix,' and the appendix section titles should be checked for consistency with the actual table and figure numbering.
Circularity Check
No circular derivation: the credible-interval claim is an empirical calibration result, and self-citations are comparative rather than load-bearing.
full rationale
The paper's derivation chain is self-contained: it specifies a hierarchical model with a Laplace working likelihood and horseshoe-family priors, derives full conditional distributions and Gibbs samplers (Appendix C), runs MCMC, and evaluates 95% empirical coverage on 1000 simulated datasets (Tables 4 and 15). No parameter is fitted to force the coverage probabilities, no coverage target enters the prior specification, and the benchmarks include external methods (HSBQR, bayesreg). The load-bearing external reference, Yang et al. (2016), is used only to state that the Laplace working likelihood needs a variance correction in low dimensions and that the correction is unavailable when p > n; the paper then tests the horseshoe priors empirically without claiming a theorem. Self-citations (Fan et al. 2024; Ren et al. 2023; Liu et al. 2024) are motivational or are used to adopt the same simulation DGP for cross-comparison, which is a legitimate comparative choice and not a reduction of the conclusion to the input. The internal inconsistency that RBRHS undercovers badly (0.649 in Table 15) while the abstract attributes valid inference to the whole horseshoe family is an overgeneralization/correctness risk, not a circular step, because the reported coverage numbers are not constructed from the conclusion. Overall, no equation in the paper reduces to its own inputs, and no prediction is forced by a fitted parameter or a self-citation chain.
Assumptions & free parameters
free parameters (3)
- sigma_beta0^2 (intercept prior variance)
- e, f (Gamma hyperparameters for tau)
- c, d (inverse-gamma hyperparameters for b^2 in RBRHS)
assumptions (3)
- standard math Laplace errors admit the scale-mixture representation epsilon_i = tau^{-1} xi sqrt(v_i) z_i with v_i ~ Exp(1), z_i ~ N(0,1), xi = sqrt(8)
- standard math A half-Cauchy prior on a scale parameter s is equivalent to s^2|nu ~ IG(1/2,1/nu), nu ~ IG(1/2,1)
- domain assumption The Laplace working likelihood yields calibrated posterior credible intervals in p > n problems once a shrinkage prior is used, even though Yang et al. (2016) showed a variance correction is required in low-dimensional Bayesian quantile regression
Cite this review
Pith. "Pith review of Robust Bayesian high-dimensional variable selection and inference with the horseshoe family of priors." pith.science (2026). https://pith.science/paper/SP7IEC25
@misc{pith2026250710975,
author = {Pith},
title = {Pith review of: Robust Bayesian high-dimensional variable selection and inference with the horseshoe family of priors},
year = {2026},
howpublished = {\url{https://pith.science/paper/SP7IEC25}},
note = {Machine review of arXiv:2507.10975}
}
read the original abstract
Frequentist robust variable selection has been extensively investigated in high-dimensional regression. Despite success, developing the corresponding statistical inference procedures remains a challenging task. Recently, tackling this challenge from a Bayesian perspective has received much attention. In literature, the two-group spike-and-slab priors that can induce exact sparsity have been demonstrated to yield valid inference in robust sparse linear models. Nevertheless, another important category of sparse priors, the horseshoe family of priors, including horseshoe, horseshoe+, and regularized horseshoe priors, has not yet been examined in robust high-dimensional regression by far. Their performance in variable selection and especially statistical inference in the presence of heavy-tailed model errors is not well understood. In this paper, we address the question by developing robust Bayesian hierarchical models utilizing the horseshoe family of priors along with an efficient Gibbs sampling scheme. We show that compared with competing methods with alternative sampling strategies such as slice sampling, our proposals lead to superior performance in variable selection, Bayesian estimation and statistical inference. In particular, our numeric studies indicate that even without imposing exact sparsity, the one-group horseshoe priors can still yield valid Bayesian credible intervals under robust high-dimensional linear regression models. Applications of the proposed and alternative methods on real data further illustrates the advantage of the proposed methods.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
A selective review of robust variable selection with applications in bioinformatics,
C. Wu and S. Ma, “A selective review of robust variable selection with applications in bioinformatics,” Briefings in bioinformatics , vol. 16, no. 5, pp. 873–883, 2015
work page 2015
-
[2]
High-dimensional inference: confidence intervals, p-values and r-software hdi,
R. Dezeure, P. B¨ uhlmann, L. Meier, and N. Meinshausen, “High-dimensional inference: confidence intervals, p-values and r-software hdi,”Statistical science, pp. 533–558, 2015
work page 2015
-
[3]
K. Fan, S. Subedi, G. Yang, X. Lu, J. Ren, and C. Wu, “Is seeing believing? a prac- titioner’s perspective on high-dimensional statistical inference in cancer genomics stud- ies,” Entropy, vol. 26, no. 9, p. 794, 2024
work page 2024
-
[4]
A review of bayesian variable selection methods: what, how and which,
R. B. O’hara and M. J. Sillanp¨ a¨ a, “A review of bayesian variable selection methods: what, how and which,” 2009
work page 2009
-
[5]
Bayesian variable selection in linear regression,
T. J. Mitchell and J. J. Beauchamp, “Bayesian variable selection in linear regression,” Journal of the american statistical association , vol. 83, no. 404, pp. 1023–1032, 1988
work page 1988
-
[6]
Variable selection via gibbs sampling,
E. I. George and R. E. McCulloch, “Variable selection via gibbs sampling,” Journal of the American Statistical Association , vol. 88, no. 423, pp. 881–889, 1993
work page 1993
-
[7]
Spike-and-slab meets lasso: A review of the spike-and-slab lasso,
R. Bai, V. Roˇ ckov´ a, and E. I. George, “Spike-and-slab meets lasso: A review of the spike-and-slab lasso,” Handbook of Bayesian variable selection , pp. 81–108, 2021
work page 2021
-
[8]
The horseshoe estimator for sparse signals,
C. M. Carvalho, N. G. Polson, and J. G. Scott, “The horseshoe estimator for sparse signals,” Biometrika, pp. 465–480, 2010
work page 2010
Show all 43 references
-
[9]
Shrink globally, act locally: Sparse bayesian regulariza- tion and prediction,
N. G. Polson and J. G. Scott, “Shrink globally, act locally: Sparse bayesian regulariza- tion and prediction,” Bayesian statistics , vol. 9, no. 501-538, p. 105, 2010
2010
-
[10]
Lasso meets horseshoe,
A. Bhadra, J. Datta, N. G. Polson, and B. Willard, “Lasso meets horseshoe,” Statistical Science, vol. 34, no. 3, pp. 405–427, 2019
2019
-
[11]
Robust bayesian variable selection for gene–environment interactions,
J. Ren, F. Zhou, X. Li, S. Ma, Y. Jiang, and C. Wu, “Robust bayesian variable selection for gene–environment interactions,” Biometrics, vol. 79, no. 2, pp. 684–694, 2023
2023
-
[12]
Uncertainty quantification for the horseshoe (with discussion),
S. van der Pas, B. Szab´ o, and A. van der Vaart, “Uncertainty quantification for the horseshoe (with discussion),” vol. 12, no. 4, pp. 1221–1274, 2017
2017
-
[13]
The horseshoe+ estimator of ultra-sparse signals,
A. Bhadra, J. Datta, N. G. Polson, and B. Willard, “The horseshoe+ estimator of ultra-sparse signals,” Bayesian Analysis, vol. 12, no. 4, pp. 1105–1131, 2017. 26
2017
-
[14]
Sparsity information and regularization in the horseshoe and other shrinkage priors,
J. Piironen and A. Vehtari, “Sparsity information and regularization in the horseshoe and other shrinkage priors,” Electronic Journal of Statistics , vol. 11, no. 2, pp. 5018– 5051, 2017
2017
-
[15]
Bayesian semiparametric modelling in quantile regression,
A. Kottas and M. Krnjaji´ c, “Bayesian semiparametric modelling in quantile regression,” Scandinavian Journal of Statistics , vol. 36, no. 2, pp. 297–319, 2009
2009
-
[16]
A general framework for updating belief distributions,
P. G. Bissiri, C. C. Holmes, and S. G. Walker, “A general framework for updating belief distributions,” Journal of the Royal Statistical Society Series B: Statistical Methodology, vol. 78, no. 5, pp. 1103–1130, 2016
2016
-
[17]
Gibbs sampling methods for bayesian quantile regres- sion,
H. Kozumi and G. Kobayashi, “Gibbs sampling methods for bayesian quantile regres- sion,” Journal of statistical computation and simulation , vol. 81, no. 11, pp. 1565–1578, 2011
2011
-
[18]
Mean field variational bayes for elaborate distributions,
M. P. Wand, J. T. Ormerod, S. A. Padoan, and R. Fr¨ uhwirth, “Mean field variational bayes for elaborate distributions,” Bayesian Analysis, vol. 6, no. 4, pp. 847–900, 2011
2011
-
[19]
A simple sampler for the horseshoe estimator,
E. Makalic and D. F. Schmidt, “A simple sampler for the horseshoe estimator,” IEEE Signal Processing Letters, vol. 23, no. 1, pp. 179–182, 2015
2015
-
[20]
The spike-and-slab quantile lasso for robust variable selection in cancer genomics studies,
Y. Liu, J. Ren, S. Ma, and C. Wu, “The spike-and-slab quantile lasso for robust variable selection in cancer genomics studies,” Statistics in Medicine , vol. 43, no. 26, pp. 4928– 4983, 2024
2024
-
[21]
The spike-and-slab lasso,
V. Roˇ ckov´ a and E. I. George, “The spike-and-slab lasso,” Journal of the American Statistical Association, vol. 113, no. 521, pp. 431–444, 2018
2018
-
[22]
Horseshoe prior bayesian quantile regression,
D. Kohns and T. Szendrei, “Horseshoe prior bayesian quantile regression,” Journal of the Royal Statistical Society Series C: Applied Statistics , vol. 73, no. 1, pp. 193–220, 2024
2024
-
[23]
Package ‘horseshoe’,
S. van der Pas, J. Scott, A. Chakraborty, A. Bhattacharya, and M. S. van der Pas, “Package ‘horseshoe’,” 2022
2022
-
[24]
High-dimensional bayesian regularised regression with the bayesreg package,
E. Makalic and D. F. Schmidt, “High-dimensional bayesian regularised regression with the bayesreg package,” arXiv preprint arXiv:1611.06649 , 2016
2016 arXiv
-
[25]
Bayesian quantile regression,
K. Yu and R. A. Moyeed, “Bayesian quantile regression,”Statistics & Probability Letters, vol. 54, no. 4, pp. 437–447, 2001. 27
2001
-
[26]
Posterior inference in bayesian quantile regression with asymmetric laplace likelihood,
Y. Yang, H. J. Wang, and X. He, “Posterior inference in bayesian quantile regression with asymmetric laplace likelihood,” International Statistical Review , vol. 84, no. 3, pp. 327–344, 2016
2016
-
[27]
Optimal predictive model selection,
M. M. Barbieri and J. O. Berger, “Optimal predictive model selection,” The Annals of Statistics, vol. 32, no. 3, pp. 870–897, 2004
2004
-
[28]
The median probability model and correlated variables,
M. M. Barbieri, J. O. Berger, E. I. George, and V. Roˇ ckov´ a, “The median probability model and correlated variables,” Bayesian Analysis, vol. 16, no. 4, pp. 1085–1112, 2021
2021
-
[29]
Bayesian estimation of sparse signals with a continuous spike-and-slab prior,
V. Roˇ ckov´ a, “Bayesian estimation of sparse signals with a continuous spike-and-slab prior,” Annals of Statistics , vol. 46, no. 1, pp. 401–437, 2018
2018
-
[30]
The bayesian elastic net,
Q. Li and N. Lin, “The bayesian elastic net,” Bayesian Analysis, vol. 5, no. 1, pp. 151– 170, 2010
2010
-
[31]
Regularization and variable selection via the elastic net,
H. Zou and T. Hastie, “Regularization and variable selection via the elastic net,”Journal of the Royal Statistical Society Series B: Statistical Methodology , vol. 67, no. 2, pp. 301– 320, 2005
2005
-
[32]
Nearly optimal bayesian shrinkage for high-dimensional regres- sion,
Q. Song and F. Liang, “Nearly optimal bayesian shrinkage for high-dimensional regres- sion,” Science China Mathematics , vol. 66, no. 2, pp. 409–442, 2023
2023
-
[33]
Inference from iterative simulation using multiple se- quences,
A. Gelman and D. B. Rubin, “Inference from iterative simulation using multiple se- quences,” Statistical science, pp. 457–472, 1992
1992
-
[34]
General methods for monitoring convergence of iterative simulations,
S. P. Brooks and A. Gelman, “General methods for monitoring convergence of iterative simulations,” Journal of computational and graphical statistics , vol. 7, no. 4, pp. 434– 455, 1998
1998
-
[35]
Gelman, J
A. Gelman, J. B. Carlin, H. S. Stern, and D. B. Rubin,Bayesian data analysis. Chapman and Hall/CRC, 1995
1995
-
[36]
Regulation of gene expression in the mammalian eye and its relevance to eye disease,
T. E. Scheetz, K.-Y. A. Kim, R. E. Swiderski, A. R. Philp, T. A. Braun, K. L. Knudtson, A. M. Dorrance, G. F. DiBona, J. Huang, T. L. Casavant, et al. , “Regulation of gene expression in the mammalian eye and its relevance to eye disease,” Proceedings of the National Academy o...
2006
-
[37]
Adaptive lasso for sparse high-dimensional regres- sion models,
J. Huang, S. Ma, and C.-H. Zhang, “Adaptive lasso for sparse high-dimensional regres- sion models,” statistica sinica, pp. 1603–1618, 2008. 28
2008
-
[38]
Quantile regression for analyzing heterogeneity in ultra- high dimension,
L. Wang, Y. Wu, and R. Li, “Quantile regression for analyzing heterogeneity in ultra- high dimension,” Journal of the American Statistical Association , vol. 107, no. 497, pp. 214–222, 2012
2012
-
[39]
Penalized generalized estimating equations for high- dimensional longitudinal data analysis,
L. Wang, J. Zhou, and A. Qu, “Penalized generalized estimating equations for high- dimensional longitudinal data analysis,” Biometrics, vol. 68, no. 2, pp. 353–360, 2012
2012
-
[40]
Penalized variable selection for lipid–environment interactions in a longitudinal lipidomics study,
F. Zhou, J. Ren, G. Li, Y. Jiang, X. Li, W. Wang, and C. Wu, “Penalized variable selection for lipid–environment interactions in a longitudinal lipidomics study,” Genes, vol. 10, no. 12, p. 1002, 2019
2019
-
[41]
Sparse group variable selection for gene–environment interactions in the longitudinal study,
F. Zhou, X. Lu, J. Ren, K. Fan, S. Ma, and C. Wu, “Sparse group variable selection for gene–environment interactions in the longitudinal study,”Genetic epidemiology, vol. 46, no. 5-6, pp. 317–340, 2022
2022
-
[42]
bayesqr: A bayesian approach to quantile regres- sion,
D. F. Benoit and D. Van den Poel, “bayesqr: A bayesian approach to quantile regres- sion,” Journal of Statistical Software , vol. 76, pp. 1–32, 2017
2017
-
[43]
Identifying gene–environment interactions with robust marginal bayesian variable selection,
X. Lu, K. Fan, J. Ren, and C. Wu, “Identifying gene–environment interactions with robust marginal bayesian variable selection,” Frontiers in Genetics , vol. 12, p. 667074, 2021. 29 A Additional simulation results A.1 Variable selection and estimation Table 8: Identification an...
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.