REVIEW 5 major objections 5 minor 55 references
Variational Garrote for Statistical Physics-based Sparse and Robust Variable Selection
T0 review · 5 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Variational Garrote beats LASSO in sparse variable selection
desk verdict Interesting uncertainty-transition heuristic, but the heterogeneous sparsity measures across methods undermine the comparative claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the variational free energy F(m,w) obtained from the VG posterior by marginalizing binary selection variables s_i with a factorized Bernoulli approximation Q(s). F combines a reconstruction energy (including a variance term from s_i^2=s_i), an entropy term that resists masks collapsing, and a -gamma sum m_i sparsity penalty; the inverse temperature beta is eliminated via beta=M/2E. Its minimizer gives masks m_i in [0,1], mean-field probabilities that variable i is selected. The paper also derives piecewise analytic mean-field estimates for selection error E_sel and selection uncertainty sigma_sel as functions of rho_model and rho_data, which let it recognize the over-selection
What would settle it
Simulate spike-and-slab data with a known number of relevant variables (say 5 of 256); grid-search the regularization parameters so that LASSO and VG admit exactly the same number of selected variables. If LASSO's selection error and generalization error then equal VG's, the claimed VG advantage is an artifact of how sparsity is measured; if VG still wins, the soft-mask mechanism is confirmed. A second check: with true rho_data known, test whether the elbow-threshold rule applied to LASSO weights recovers rho_model = rho_data at optimal regularization; a systematic mismatch would invalidate th
Extended reading notes
Core claim
On synthetic data whose true weights are drawn from a spike-and-slab distribution with known density rho_data, the paper compares Ridge, LASSO, and VG at matched model density rho_model. Its central finding is that in the highly sparse regime (rho_data around 2% relevant variables), VG has the lowest generalization error and the lowest selection error E_sel. The reason the paper identifies is the soft mask: VG assigns fractional mask values 0<m_i<1 to several true predictors, whereas the convex losses of Ridge and LASSO tend to concentrate on one dominant variable, missing equally relevant alternatives. As rho_model grows past rho_data, generalization degrades sharply and the ensemble select
Load-bearing premise
The comparison treats each method's inferred density rho_model as the same measure of sparsity, but for Ridge and LASSO that density comes from ad hoc thresholds; if those thresholds mis-calibrate what 'equally sparse' means, the ranking of the methods is not established.
Editorial extensions
If this is right
- In highly sparse, underdetermined regression, VG can recover a small relevant subset more consistently than Ridge or LASSO at the same model density.
- The minimum of selection error E_sel occurs near rho_model = rho_data for all three methods, supporting rho_model as a meaningful sparsity axis and predicting that model selection should target the density of the data.
- The sharp increase in selection uncertainty when too many variables are allowed gives a stopping rule for adding variables; in real data it yielded 2.5-4 relevant variables for crime prediction and 1 for blog feedback.
- Because VG's masks are soft rather than hard zeros, the method can keep several candidate variables alive simultaneously, avoiding the single-dominant-variable bias of convex methods in sparse regimes.
- Extending VG with automatic differentiation makes the variational loss trainable with modern optimizers, so the method scales beyond the small problems of the original formulation.
- pith_inferences
Reading between the lines
- A natural next test, not pursued here, is whether the sharp sigma_sel transition persists with correlated inputs and heteroscedastic noise; the synthetic experiments use independent Gaussian predictors, while real data have correlations, so the observed transition in the real datasets does not yet isolate the mechanism.
- The mean-field inversion for rho_data could be applied to any model that outputs selection probabilities, not only VG, turning the uncertainty curve into a general sparsity-estimation diagnostic.
- The paper's poor VG performance in the high-density regime suggests a practical hybrid: use Ridge or LASSO as a dense prescreener, then run VG on the surviving variables for final sparse selection.
- A direct oracle-sparsity comparison, where all methods are forced to the same support size, would separate the method's selection quality from the inconvenience that LASSO and Ridge sparsity must be inferred from weights.
- supporting_citations
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper revisits the Variational Garrote (VG), a Bayesian variable-selection method with explicit binary selection variables, and proposes a modern automatic-differentiation implementation. The authors compare VG against Ridge and LASSO on synthetic spike-and-slab data (N=M=256) at three true sparsity levels and on two UCI regression datasets. They report that VG is especially accurate in highly sparse regimes, shows more consistent selection across sparsity levels, and exhibits a sharp transition in selection uncertainty that can be used to estimate the number of relevant variables. A mean-field theory is presented for selection error and selection uncertainty, and the latter is inverted on real data to infer data sparsity.
Significance. If the claims hold, the work provides a practical, scalable implementation of a relatively underused sparse-regression method and offers an interpretable, analytically motivated uncertainty curve for choosing the number of relevant variables. The mean-field derivation of E_sel and sigma_sel is transparent, and the use of 20,000 synthetic ensembles is a strength. However, the central comparative claims and the proposed transition signal rest on a sparsity measure that is defined differently across methods and on a theoretical curve whose assumptions are not tested. Without error bars, a valid teacher distribution, and a demonstrated real-data generalization improvement, the significance is limited; the conclusions are currently not fully supported by the evidence presented.
major comments (5)
- [Sec. III.A.2 and III.B] The central comparison uses rho_model as a common sparsity measure, but its definition is model-dependent. For VG it is the sum of soft masks; for LASSO it is a count above a zero-centered binary-mixture 'elbow'; for Ridge it is a count above a threshold derived from the smallest weights of an unregularized fit. No evidence is given that equal rho_model corresponds to equal model complexity. A concrete symptom: inverting Eq. (19) on the Community Crimes data yields rho_data estimates of 3 (Ridge), 4 (LASSO), and 2.5 (VG). If the sparsity measures were comparable and the theory correct, these estimates should agree. Their disagreement undermines both the comparative performance claim and the practical use of the transition curve.
- [Eq. (14)] The spike-and-slab teacher distribution is not normalized for the densities used. The slab density is written as (1/2) rho_data on each side over 1<|w|<bar{w}, so the total slab mass is rho_data(bar{w}-1). With bar{w}=sqrt(12/rho_data-3/4)-1/2, for rho_data=80/256 one obtains bar{w}≈5.64 and slab mass ≈1.45, giving total probability >2. Thus the actual fraction of nonzero teacher weights is not rho_data for the intermediate and high-density regimes, which are precisely the regimes in Figs. 2(a,b) and 3(d,f). This invalidates the controlled sparsity calibration on which the comparison relies.
- [Figs. 2 and 3] All performance and selection curves are ensemble averages over 20,000 realizations, but no error bars or confidence intervals are shown. Claims such as 'VG achieves the lowest E_gen' in Fig. 2(c) and 'VG shows a sharp transition in sigma_sel' cannot be assessed for statistical significance. Since the practical proposal (use the transition to estimate the true number of variables) depends on the sharpness and reliability of the increase in sigma_sel, a quantification of ensemble variability is necessary.
- [Sec. IV and Fig. 4] The conclusion states that VG 'achieves low generalization error' on real-world data, but Fig. 4 reports only masks, sigma_sel, and inferred rho_data distributions. No real-data E_gen results are shown for CC or BF. Additionally, the 'sharp transition' is not directly observed on real data; it is assumed from Eq. (19) and then used to infer rho_data. This is partly circular and does not validate the proposed signal as a general tool.
- [Sec. III.A.4] The paper acknowledges that for LASSO and VG it is difficult to tune rho_model below rho_data, and that the low-density visualizations are based on extrapolation. This is a load-bearing limitation: the claimed advantage of VG in the under-selection regime rests on this extrapolated region, and the experimental basis for the left branch of the V-shaped curves is thin.
minor comments (5)
- [Eq. (18)] The definition of sigma_sel uses ⟨m_i⟨⟨1−m_i⟩, which is ambiguous. The appendix and subsequent usage suggest ⟨m_i⟩(1−⟨m_i⟩), but the displayed formula should be made explicit.
- [Appendix B] Calling the normalized mixture coefficients P(rho_data) is misleading: they are coefficients of a nonnegative linear regression, not a Bayesian posterior. Please rename or explain the interpretation.
- [References] Reference [40] is incomplete: 'arXiv: Statistics Theory (2010)' lacks a title and ID. The RSS discussion of LASSO [26] should also cite Tibshirani's original paper correctly.
- [Sec. III.A.2] The statement that for Ridge 'all mask values are theoretically m_i=1' is confusing because Ridge weights are nonzero but are not selection masks; please clarify that this refers to the absence of exact zeros.
- [Sec. II.B] The derivation of Eq. (13) eliminates beta using the stationarity condition (12), but the text does not note that this makes the loss a profile likelihood rather than the original variational bound. This is a useful clarification for readers.
Circularity Check
No significant circularity; the formal derivations are self-contained, though cross-model sparsity comparability is a methodological concern.
full rationale
The paper's central formal result, the mean-field expression for the selection uncertainty sigma_sel in Eq. (19), is derived explicitly in Appendix A from stated assumptions (no false positives in under-selection, no false negatives in over-selection, and uniform mask assignment among the selected/relevant variables). This is a mathematical derivation, not a fit to the observed sigma_sel, and its agreement with the synthetic-data points in Figs. 3(e,f) is an empirical check rather than a tautology. The real-data inference of rho_data by fitting Eq. (19) to observed sigma_sel is a standard model-inversion procedure, not a prediction that reduces to its input. No load-bearing self-citations appear: the VG objective is attributed to Kappen and Gomez [30], not to the present authors, and no uniqueness theorem is imported from the authors' prior work. The main weakness is that rho_model is measured differently across methods—soft masks for VG, an elbow/binary-mixture threshold for LASSO, and a threshold derived from an unregularized fit for Ridge—which threatens the comparability of the x-axis in cross-model comparisons. However, that is a measurement-validity issue, not a circularity. No equation in the paper is equivalent to its inputs by construction, and no fitted parameter is renamed as a prediction. Therefore, the paper is substantially self-contained with respect to its formal derivations.
Assumptions & free parameters
free parameters (4)
- gamma (VG sparsity prior) =
varied to realize rho_model values; exact grid not reported
- LASSO elbow threshold =
fit per dataset via zero-centered binary mixture
- Ridge lower bound w* =
average of smallest absolute weights from unregularized linear regression
- Regularization strengths lambda for LASSO and Ridge =
varied to sweep rho_model
assumptions (5)
- domain assumption Linear regression with additive Gaussian noise and independent flat priors for VG weights and beta.
- domain assumption Teacher weights are drawn from a spike-and-slab distribution with variance normalized to 2 and slab support |w| in [1, wbar].
- ad hoc to paper Mean-field approximation in Appendix A: under-selection has no false positives and all relevant masks are equal; over-selection has no false negatives and all irrelevant masks are equal.
- ad hoc to paper Observed sigma_sel can be represented as a nonnegative linear combination of single-density template curves f(rho_model; rho_data).
- domain assumption The processed UCI datasets have approximately sparse linear structure relevant to the comparison.
Cite this review
Pith. "Pith review of Variational Garrote for Statistical Physics-based Sparse and Robust Variable Selection." pith.science (2026). https://pith.science/paper/BZCGONFJ
@misc{pith2026250906383,
author = {Pith},
title = {Pith review of: Variational Garrote for Statistical Physics-based Sparse and Robust Variable Selection},
year = {2026},
howpublished = {\url{https://pith.science/paper/BZCGONFJ}},
note = {Machine review of arXiv:2509.06383}
}
read the original abstract
Selecting key variables from high-dimensional data is increasingly important in the era of big data. Sparse regression serves as a powerful tool for this purpose by promoting model simplicity and explainability. In this work, we revisit a valuable yet underutilized method, the statistical physics-based Variational Garrote (VG), which introduces explicit feature selection spin variables and leverages variational inference to derive a tractable loss function. We enhance VG by incorporating modern automatic differentiation techniques, enabling scalable and efficient optimization. We evaluate VG on both fully controllable synthetic datasets and complex real-world datasets. Our results demonstrate that VG performs especially well in highly sparse regimes, offering more consistent and robust variable selection than Ridge and LASSO regression across varying levels of sparsity. We also uncover a sharp transition: as superfluous variables are admitted, generalization degrades abruptly and the uncertainty of the selection variables increases. This transition point provides a practical signal for estimating the correct number of relevant variables, an insight we successfully apply to identify key predictors in real-world data. We expect that VG offers strong potential for sparse modeling across a wide range of applications, including compressed sensing and model pruning in machine learning.
Figures
Reference graph
Works this paper leans on
-
[1]
Experimental setup We begin by examining sparse regression using a syn- thetic dataset that allows for systematic control. This dataset is designed to emulate real-world conditions, where only a small subset of variables has a significant impact on the target. Since regression performance can depend on the true distribution of regression weights, we adopt...
-
[2]
Variable sparsity To enable a fair comparison of sparse models, we face a challenge: the regularization terms differ in form, each with its own sparsity-controlling parameters. To address this, we introduce a unified measure for variable sparsity that can be applied consistently across all three models of Ridge, LASSO, and VG. In the VG model, sparsity ca...
-
[3]
Generalization performance We now evaluate the performance of sparse regression models under varying numbers of relevant variables. The primary performance metric is the generalization error, which quantifies the model’s ability to predictyfor previ- ously unseen input datax. It is defined as the normalized distance between the estimated valuesy µ and tru...
-
[4]
Variable selection performance We next examine how each sparse model’s variable se- lection evolves as the permitted densityρ model increases. By focusing on low-density regimes, we can assess each model’s capability to accurately identify the true relevant variables. Figures 3(a) and (b) visually illustrate how the three sparse regression models progress...
-
[5]
S. Chatterjee and A. S. Hadi,Regression analysis by ex- ample(John Wiley & Sons, 2015)
work page 2015
- [6]
-
[7]
A. J. Majda and J. Harlim, Nonlinearity26, 201 (2012)
work page 2012
-
[8]
A. A. Kaptanoglu, C. Hansen, J. D. Lore, M. Landreman, and S. L. Brunton, Phys. Plasmas30(2023)
work page 2023
Show all 55 references
-
[9]
S. L. Brunton, J. L. Proctor, and J. N. Kutz, Proc. Natl. Acad. Sci. U.S.A.113, 3932 (2016). 11
2016
-
[10]
Quade, M
M. Quade, M. Abel, K. Shafi, R. K. Niven, and B. R. Noack, Phys. Rev. E94, 012214 (2016)
2016
-
[11]
Maalouf, International Journal of Data Analysis Tech- niques and Strategies3, 281 (2011)
M. Maalouf, International Journal of Data Analysis Tech- niques and Strategies3, 281 (2011)
2011
-
[12]
Zhang, inProceedings of the twenty-first international conference on Machine learning(2004) p
T. Zhang, inProceedings of the twenty-first international conference on Machine learning(2004) p. 116
2004
-
[13]
R. M. Dawes and B. Corrigan, Psychol. Bull.81, 95 (1974)
1974
-
[14]
D. C. Montgomery, E. A. Peck, and G. G. Vining,Intro- duction to linear regression analysis(John Wiley & Sons, 2021)
2021
-
[15]
Cramer, Tinbergen Institute, Tinbergen Institute Dis- cussion Papers (2002), 10.2139/ssrn.360300
J. Cramer, Tinbergen Institute, Tinbergen Institute Dis- cussion Papers (2002), 10.2139/ssrn.360300
2002 doi
-
[16]
Cortes and V
C. Cortes and V. Vapnik, Mach. Learn.20, 273 (1995)
1995
-
[17]
J. A. Nelder and R. W. M. Wedderburn, J. R. Stat. Soc. A135, 370 (1972)
1972
-
[18]
W. Zhou, Z. Yan, and L. Zhang, Sci. Rep.14, 5905 (2024)
2024
-
[19]
B. S. Shastry, J. Hum. Genet.52, 871 (2007)
2007
-
[20]
Olsson, R
B. Olsson, R. Lautner, U. Andreasson, A. ¨Ohrfelt, E. Portelius, M. Bjerke, M. H¨ oltt¨ a, C. Ros´ en, C. Olsson, G. Strobel, E. Wu, K. Dakin, M. Petzold, K. Blennow, and H. Zetterberg, Lancet Neurol.15, 673 (2016)
2016
-
[21]
Einav and J
L. Einav and J. Levin, Science346, 1243089 (2014)
2014
-
[22]
S. Han, J. Pool, J. Tran, and W. J. Dally, inProceed- ings of the 29th International Conference on Neural In- formation Processing Systems - Volume 1, NIPS’15 (MIT Press, Cambridge, MA, USA, 2015) p. 1135–1143
2015
-
[23]
Frankle and M
J. Frankle and M. Carbin, inInternational Conference on Learning Representations(2019)
2019
-
[24]
Clark, U
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning, inProceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, edited by T. Linzen, G. Chrupa la, Y. Belinkov, and D. Hupkes (Association for Computational Linguistics, Florence, Ital...
2019
-
[25]
Michel, O
P. Michel, O. Levy, and G. Neubig, inAdvances in Neu- ral Information Processing Systems, Vol. 32, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch´ e- Buc, E. Fox, and R. Garnett (Curran Associates, Inc., 2019)
2019
-
[26]
M. A. Babyak, Biopsychosocial Science and Medicine66, 411 (2004)
2004
-
[27]
D. M. Hawkins, J. Chem. Inf. Comput. Sci.44, 1 (2004)
2004
-
[28]
Gruber,Improving Efficiency by Shrinkage: The James–Stein and Ridge Regression Estimators, 1st ed
M. Gruber,Improving Efficiency by Shrinkage: The James–Stein and Ridge Regression Estimators, 1st ed. (Routledge, 1998)
1998
-
[29]
A. E. Hoerl and R. W. Kennard, Technometrics12, 55 (1970)
1970
-
[30]
Tibshirani, J
R. Tibshirani, J. R. Stat. Soc. B (Methodol.)58, 267 (1996)
1996
-
[31]
Sparse autoencoder,
A. Ng, “Sparse autoencoder,” (2011), unpublished lec- ture notes
2011
-
[32]
Huben, H
R. Huben, H. Cunningham, L. R. Smith, A. Ewart, and L. Sharkey, inThe Twelfth International Conference on Learning Representations(2024)
2024
-
[33]
Breiman, Technometrics37, 373 (1995)
L. Breiman, Technometrics37, 373 (1995)
1995
-
[34]
H. J. Kappen and V. G´ omez, Mach. Learn.96, 269 (2014)
2014
-
[35]
Penrose, Math
R. Penrose, Math. Proc. Cambridge Philos. Soc.51, 406–413 (1955)
1955
-
[36]
Beck and M
A. Beck and M. Teboulle, SIAM J. Imaging Sci.2, 183 (2009)
2009
-
[37]
Friedman, T
J. Friedman, T. Hastie, and R. Tibshirani, J. Stat. Softw. 33, 1 (2010)
2010
-
[38]
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein, Found. Trends Mach. Learn.3, 1 (2011)
2011
-
[39]
Zou and T
H. Zou and T. Hastie, J. R. Stat. Soc. B67, 301 (2005)
2005
-
[40]
T. J. Mitchell and J. J. Beauchamp, J. Am. Stat. Assoc. 83, 1023 (1988)
1988
-
[41]
E. I. George and R. E. McCulloch, Stat. Sin.7, 339 (1997)
1997
-
[42]
M´ ezard, G
M. M´ ezard, G. Parisi, and M. A. Virasoro,Spin Glass Theory and Beyond(World Scientific, 1987)
1987
-
[43]
Zhang and J
C.-H. Zhang and J. Huang, The Annals of Statistics36, 1567 (2008)
2008
-
[44]
Zhou, arXiv: Statistics Theory (2010)
S. Zhou, arXiv: Statistics Theory (2010)
2010
-
[45]
Communities and Crime,
M. Redmond, “Communities and Crime,” UCI Machine Learning Repository (2002), DOI: https://doi.org/10.24432/C53W3X
2002 doi
-
[46]
BlogFeedback,
K. Buza, “BlogFeedback,” UCI Machine Learning Repos- itory (2014), DOI: https://doi.org/10.24432/C58S3F
2014 doi
-
[47]
S. T. Hansen, C. Stahlhut, and L. K. Hansen, inProc. 2013 Int. Workshop on Pattern Recognition in Neu- roimaging (PRNI)(IEEE, 2013) pp. 106–109
2013
-
[48]
M. R. Andersen, S. T. Hansen, and L. K. Hansen, in 2013 IEEE Int. Workshop on Machine Learning for Sig- nal Processing (MLSP)(IEEE, 2013) pp. 1–6
2013
-
[49]
D. L. Donoho, IEEE Trans. Inf. Theory52, 1289 (2006)
2006
-
[50]
Lustig, D
M. Lustig, D. Donoho, and J. M. Pauly, Magn. Reson. Med.58, 1182 (2007)
2007
-
[51]
C. H. Martin and M. W. Mahoney, J. Mach. Learn. Res. 22, 1 (2021)
2021
-
[52]
C. H. Martin and M. W. Mahoney, inProceedings of the 2020 SIAM International Conference on Data Mining (SDM)(2020) pp. 505–513
2020
-
[53]
Meng and J
X. Meng and J. Yao, J. Mach. Learn. Res.24, 1 (2023)
2023
-
[54]
Louizos, K
C. Louizos, K. Ullrich, and M. Welling, inAdvances in Neural Information Processing Systems (NeurIPS) (2017)
2017
-
[55]
Louizos, M
C. Louizos, M. Welling, and D. P. Kingma, inInterna- tional Conference on Learning Representations (ICLR) (2018)
2018
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.