Pith. sign in

REVIEW 2 major objections 3 minor 43 references

From Graphical Lasso to Atomic Norms: High-Dimensional Pattern Recovery

T0 review · 2 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper proves that for atomic-norm penalties whose unit ball is a polytope, exact pattern recovery in precision matrix estimation holds whenever a generalized irrepresentability condition is met and the true pattern's face projection…

desk verdict Genuine generalization of GLASSO pattern recovery to atomic norms; the headline ℓ1 improvement is real, but the abstract oversells the irrepresentability gain and the R2011 correction needs a referee's check. read the letter →

arxiv 2506.13353 v2 pith:G6PTGVI3 submitted 2025-06-16 math.ST stat.TH

classification math.STstat.TH MSC 62F1262H22
keywords precisionmatrixestimationatomicnormpatternrecoverygraphicallassocoloredmodelsirrepresentabilityconditionSLOPEpolytopefacialstructure
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper extends high-dimensional pattern recovery for precision matrix estimation from the graphical lasso's ℓ1 penalty to any atomic norm whose unit ball is a polytope. It establishes a theorem: under a generalized irrepresentability condition and a positive threshold τ⋄, if the sample covariance is within a deviation δ of the true covariance, then the penalized estimator recovers exactly the same pattern as the true precision matrix. The pattern is understood through the facial structure of the dual unit ball, so sparsity, equality constraints, clusters and sign patterns are all treated by one mechanism. Specialized to the ℓ1 penalty, the result gives weaker deviation requirements and better asymptotics than prior work, and the paper identifies a proof error in that prior work's main theorem.

What carries the argument

The machinery is the facial geometry of the dual unit ball B∗ of the atomic norm. Patterns of a vector x are identified with the subdifferential ∂∥x∥⋄, a face F_I of B∗; the pattern subspace S_I is the linear span of the pattern class. The face projection f_I=P_I v_{i0} summarizes the active face. Two thresholds carry the argument: τ⋄(I) is the largest tangential perturbation of f_I that keeps the shifted point inside B∗, and ζ⋄(K∗) is the distance (in Γ∗-weighted norm) to the closest matrix in the pattern subspace with a different pattern. The generalized irrepresentability condition (3.1) bounds the interaction between the off-model component of Γ∗ and f_I by (1−α)τ⋄(I∗). A primal-dual witness construction, residual bound via Γ∗-weighted norms, and Brouwer fixed-point argument convert these into the deviation threshold δ.

What would settle it

Construct the four-vertex atomic norm of Example 2.6 with parameters on the boundary where τ⋄=0, e.g. α=1 or $α^{2}$+$β^{2}$=|α|, choose a true precision matrix whose active face is one of the nontrivial faces, and let n grow with λ_n→0 and √nλ_n→∞ while the irrepresentability condition holds. If the empirical pattern recovery probability tends to 1 rather than remaining bounded away from 1, the paper's central claim about the necessity of a positive τ⋄ would be wrong.

Watch

Extended reading notes

Core claim

The central claim is Theorem 3.4: assume τ⋄(I∗)>0 and that the generalized irrepresentability condition (3.1) holds with slack α∈(0,1). Then there is an explicit δ>0 such that whenever ||vec(Σˆ−Σ∗)||<δ, the unique minimizer of the log-likelihood with atomic-norm penalty recovers the true pattern, patt⋄(Kˆ)=patt⋄(K∗), and satisfies ||Γ∗vec(Kˆ−K∗)||≤(1−1/√(1+M))/η. The threshold δ is positive and has closed-form expansions; optimization over r and λ in the general theorem yields this sharpened form. This is a direct generalization of the primal-dual witness analysis of Ravikumar et al. (2011), with a reworked residual control that yields tighter bounds and, in the ℓ1 case, an asymptotic improvement of order $p^{3}$ over the earlier result. The paper also states that the earlier proof contains a small error requiring an extra factor in the deviation bound, which affects downstream sample size statements.

Load-bearing premise

The argument requires that the true pattern's face projection f_{I*} lie in the relative interior of its face of the dual unit ball, so the threshold τ⋄(I*) is strictly positive; if it sits on the boundary, the proof breaks down and Appendix C shows pattern recovery itself fails for skewed gauges.

Editorial extensions

If this is right

  • For ℓ1-penalized GLASSO, sign recovery holds under a deviation bound δ that is asymptotically up to a factor of p^3 larger than the corrected Ravikumar et al. bound in Example 3.6.
  • The previously stated sample-size exponent in Ravikumar et al. (2011) should be corrected; downstream statements such as Wainwright (2019) Proposition 11.10 would require a fourth power of (1+8/α) rather than the square.
  • For any polytope atomic norm, recoverable patterns include sparsity, equality constraints, clusters of equal-magnitude entries and hierarchical SLOPE patterns, so colored graphical model estimation falls under the same theorem.
  • Finite-sample versions follow by combining δ with existing covariance concentration inequalities, since δ itself does not depend on n.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The threshold τ⋄ could serve as a design criterion for choosing penalty weights: Appendix D already maximizes δ over SLOPE weights for known pattern classes, and the same logic extends to other atomic norms.
  • The generalized irrepresentability condition is stated for precision matrix estimation, but the paper's framework (patterns as faces, thresholds) transfers to linear regression and other M-estimators; if the transfer holds, the same τ⋄>0 condition would mark the boundary of model selection consistency there.
  • For skewed gauges with τ⋄=0, the failure of recovery is structural rather than a proof artifact: with λ_n→0 and √nλ_n→∞, the normalized subgradient stays outside the dual ball, suggesting no sample size can rescue pattern recovery without changing the penalty.
  • The numerical gap between theoretical and empirical thresholds for SLOPE suggests the ℓ∞ norm used to measure deviations is not the right metric for structured patterns; a norm adapted to the pattern subspace could tighten the bounds substantially.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper extends the primal-dual witness analysis of the graphical lasso to general polyhedral atomic norm penalties for precision matrix estimation. Patterns are identified with faces of the dual unit ball, and the paper introduces a first threshold tau_diamond controlling perturbations orthogonal to the true pattern face and a second threshold zeta_diamond measuring pattern stability along the pattern subspace. Under a generalized irrepresentability condition and the assumption tau_diamond(I*)>0, Theorems 3.3 and 3.4 give deterministic deviation bounds on ||vec(Sigma_hat - Sigma*)|| under which the penalized estimator recovers the true pattern. The results are specialized to the l1 penalty in Theorem 3.5, where the paper claims weaker deviation requirements and a factor p^3 asymptotic improvement over Ravikumar et al. (2011), supported by Example 3.6 and by numerical comparisons in Section 4.

Significance. If the results hold, this is a valuable extension of the graphical lasso theory to a broad class of structured penalties relevant to colored graphical models, and the tau_diamond/zeta_diamond framework gives a unified geometric perspective on irrepresentability conditions. The paper is honest about the main limitation: tau_diamond(I*)>0 is equivalent to the face projection lying in the relative interior of the corresponding face, and Appendix C shows that pattern recovery can fail when this fails. The appendix proofs are detailed, and the numerical experiments show genuine tightness of the l1 threshold. The main quantitative comparison with prior work, however, rests on a corrected restatement of Ravikumar et al. (2011) and on Example 3.6, where one of the displayed constants is inconsistent with its own definition; these issues are fixable and do not appear to change the asymptotic order of the claimed improvement.

major comments (2)
  1. [Section 3.4, Example 3.6] The displayed formula for kappa_Sigma* is inconsistent with its definition kappa_Sigma* = |||Sigma*|||_infinity. For rho>0, Sigma* has diagonal entries (1+(p-3)rho)/((1-rho)(1+(p-2)rho)) and off-diagonal entries -rho/((1-rho)(1+(p-2)rho)) inside the first p-1 block, so the maximum absolute row sum is (1+(2p-5)rho)/((1-rho)(1+(p-2)rho)), not (1+(p-3)rho)/((1-rho)(1+(p-2)rho)) as stated. Since kappa_Sigma* enters delta_R through kappa_Sigma*^3, the numerical values in Table 2 for the dense graph and the displayed ratio in Example 3.6 should be recomputed. I stress that the value kappa_Gamma* = (1+(p-2)rho)^2 is consistent with the support definition used in the paper, because S includes diagonal entries and both orientations of each off-diagonal pair, so Gamma*_{S,S} = Sigma_block tensor Sigma_block and its inverse has row sum (1+(p-2)rho)^2; the inconsistency is specifically in kappa_Sigma*. The asymptotic order p^3 is unaffected because the corrected kappa_Sigma* is still O(1), but the quantitative comparison needs correction.
  2. [Section 3.4, restatement of Ravikumar et al. (2011)] The paper claims that the original proof of Theorem 1 in Ravikumar et al. contains a small error and that the correct threshold is given by (3.8) with an extra factor (1+8/alpha) in the denominator compared with (3.5). This corrected delta_R is the baseline used in Table 2 and in the asymptotic comparison, so the claimed 'improvement over prior work' depends on this correction. Please provide the exact equation and page in Ravikumar et al. and a line-by-line verification that their displayed bound indeed contains the factor (1+8/alpha)^2. If the published bound is as in (3.5), then delta_R should be larger by a factor (1+8/alpha), which would reduce the reported ratio by that factor; the p^3 asymptotic order would survive, but the numerical comparison and the statement 'several orders of magnitude' would need to be adjusted accordingly.
minor comments (3)
  1. [Appendix A, Lemma 2.10] The proof of Lemma 2.10 is only a geometric sketch. Since this lemma is used to interpret the central hypothesis tau_diamond(I*)>0 and to justify the statement 'tau_diamond(I*)>0 iff f_I* is in ri(F_I*)', please expand it into a rigorous convex-analysis argument, for example by using the fact that the intersection of the dual ball with the affine hull of a face is the face itself.
  2. [Appendix D] Lemmas D.1 and D.2 are stated without proof. They are used in the SLOPE tuning discussion and in the numerical experiments, so even if they are ancillary to the main theorems, the paper should either provide proofs or give a precise reference where these calculations appear.
  3. [Section 4, Table 2] Table 2 should state explicitly which definition of delta_R is used, namely the corrected one in (3.8) rather than the published one in (3.5), and should report the value of alpha used for each graph. This will make the comparison reproducible and will prevent a reader from attributing the improvement to the original Ravikumar et al. statement.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity; the main theorems are proven and the ℓ1 comparison is benchmarked externally.

full rationale

The paper's main derivation is self-contained: Theorems 3.3 and 3.4 establish pattern recovery from a deviation bound on ||vec(hat Sigma - Sigma*)|| via the primal-dual witness method, residual control lemmas, and a Brouwer fixed-point argument (Lemmas 5.1-5.6). The conclusion patt_diamond(hat K) = patt_diamond(K*) follows from the independently proven closeness bound ||Gamma* vec(hat K - K*)|| <= r/eta together with the geometric stability radius zeta_diamond(K*), which is a condition on the true K*, not a fitted parameter. The l1 specialization (Theorem 3.5) is benchmarked against Ravikumar et al. (2011), an external result; the paper's proposed correction to delta_R is a proof claim about that paper and does not rely on the present authors' own theorems. Self-citations to Graczyk et al. (2023) supply facial-geometry definitions of patterns and to Bogdan et al. (2022) supply SLOPE pattern representations; these are prior published results with stated assumptions that do not include the paper's target, and they are not used as a uniqueness argument to exclude alternatives. Appendix C openly discusses a regime (tau_diamond = 0) where pattern recovery fails, which is an honestly disclosed limitation rather than a circular rescue. The skeptical concern about Example 3.6's claimed value kappa_Gamma* = (1+(p-2)rho)^2 is a quantitative correctness risk in the comparison with Ravikumar et al.; if the direct calculation is wrong, the advertised p^3 improvement would be weakened, but this is not a circularity because the value is not fitted from the target conclusion and is not equivalent to the paper's input by construction. Similarly, Lemma D.1/D.2 are stated without proof, which is an omitted-proof verification gap, not a circular step.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The main theorems introduce no data-fitted parameters; the constants α, η, c⋄, and the threshold values τ⋄, ζ⋄ are deterministic features of the model and norm. The assumptions are structural: positive definiteness, polytope unit ball, the irrepresentability condition, and the τ⋄>0 (face-projection interior) condition. No new physical or probabilistic entities are postulated.

assumptions (5)
  • domain assumption True covariance Σ* is positive definite and the precision matrix K*=(Σ*)^-1 exists.
    Assumed in Section 1.1; the log-determinant objective requires invertibility, and the analysis works in Sym+(p).
  • domain assumption The atomic norm unit ball is a polytope, so ∥x∥⋄ = max_i v_i^T x for a finite set {v_i}.
    Section 2.1 restricts to polyhedral gauges; patterns are defined through faces of the dual polytope.
  • domain assumption The generalized irrepresentability condition (3.1) holds for some α∈(0,1).
    Definition 3.1; analogue of the standard irrepresentability condition, needed for the dual certificate construction.
  • ad hoc to paper The face projection lies in the relative interior: fI*∈ri(FI*), so τ⋄(I*)>0.
    Theorem 3.3 requires τ⋄>0; Lemma 2.10 equates this to fI*∈ri(FI*). Appendix C shows the negative case fails, making this a structural assumption tailored to the proof technique.
  • standard math Standard convex analysis, Moore-Penrose inverses, and Brouwer's fixed point theorem.
    Used throughout Section 5 and the appendix for the fixed-point construction and subdifferential characterizations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Graphical Lasso to Atomic Norms: High-Dimensional Pattern Recovery." pith.science (2026). https://pith.science/paper/G6PTGVI3

@misc{pith2026250613353,
  author       = {Pith},
  title        = {Pith review of: From Graphical Lasso to Atomic Norms: High-Dimensional Pattern Recovery},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/G6PTGVI3}},
  note         = {Machine review of arXiv:2506.13353}
}
abstract

Estimating high-dimensional precision matrices is a fundamental problem in modern statistics, with the graphical lasso and its $\ell_1$-penalty being a standard approach for recovering sparsity patterns. However, many statistical models, e.g. colored graphical models, exhibit richer structures like symmetry or equality constraints, which the $\ell_1$-norm cannot adequately capture. This paper addresses the gap by extending the high-dimensional analysis of pattern recovery to a general class of atomic norm penalties, particularly those whose unit balls are polytopes, where patterns correspond to the polytope's facial structure. We establish theoretical guarantees for recovering the true pattern induced by these general atomic norms in precision matrix estimation. Our framework builds upon and refines the primal-dual witness methodology of Ravikumar et al. (2011). Our analysis provides conditions on the deviation between sample and true covariance matrices for successful pattern recovery, given a novel, generalized irrepresentability condition applicable to any atomic norm. When specialized to the $\ell_1$-penalty, our results offer improved conditions -- including weaker deviation requirements and a less restrictive irrepresentability condition -- leading to tighter bounds and better asymptotic performance than prior work. The proposed general irrepresentability condition, based on a new thresholding concept, provides a unified perspective on model selection consistency. Numerical examples demonstrate the tightness of the derived theoretical bounds.

Figures

Figures reproduced from arXiv: 2506.13353 by the authors.

Figure 1
Figure 1. Solid parallelogram: the dual ball B∗ , dashed parallelogram: the ball B, blue dotted line: S{1,4} = S{2,3}, red dots: face projections. Note that S{1,2} = S{3,4} is the x-axis. Left: (α, β) = (0.5, 0.8). Middle: (α, β) = (2, 1). Right: (α, β) = (0.5, 0.3). The figures are rescaled. 2.4. Vectorization. Let p ∈ N. Let m = p(p − 1)/2. Let Sym(p) denote the space of symmetric real p×p matrices, and let Sym+(p) be the c… view at source ↗
Figure 2
Figure 2. Consider a SLOPE norm with w = (1, 1/2) and a pattern I = (−1, −1). The thick black line corresponds to SI = R(1, 1)⊤. The red dot represents the face projection fI . Remark 2.8. Recall that in general we have 1 = ∥fI∥ ∗ ⋄|I ≤ ∥fI∥ ∗ ⋄ . The threshold τ⋄(I) is well-defined if and only if the set {t ≥ 0: ∀ π ⊥ ∈ S⊥ I [[Dπ⊥]] ≤ t =⇒ [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. GLASSO: Scatter plot of ∥vec(Σˆ − Σ ∗ )∥∞ (x-axis) versus ∥vec(Kˆ − K∗ )∥∞ (y-axis) for p = 25. The gray dashed line marks the theoretical threshold δ; the orange dashed line shows its empirical thresh￾old ˆδ. Points are colored blue when the true pattern (sign) is recovered and red otherwise. We carried out an analogous evaluation for the SLOPE penalty, using the weights specified in (4.1), and display the outcomes… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: SLOPE: Scatter plot of ∥vec(Σˆ − Σ ∗ )∥∞ (x-axis) versus ∥vec(Kˆ − K∗ )∥∞ (y-axis) for p = 25. The gray dashed line marks the theoretical threshold δ; the orange dashed line shows its empirical thresh￾old ˆδ. Points are colored blue when the true pattern is recovered a…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

43 extracted references · 23 canonical work pages

  1. [1]

    Litvak, Alain Pajor, and Nicole Tomczak-Jaegermann

    Rados aw Adamczak, Alexander E. Litvak, Alain Pajor, and Nicole Tomczak-Jaegermann. Quantitative estimates of the convergence of the empirical covariance matrix in log-concave ensembles. J. Amer. Math. Soc., 23 0 (2): 0 535--561, 2010. ISSN 0894-0347. doi:10.1090/S0894-0347-09-00650-X. URL https://doi.org/10.1090/S0894-0347-09-00650-X

  2. [2]

    Model selection through sparse maximum likelihood estimation for multivariate G aussian or binary data

    Onureena Banerjee, Laurent El Ghaoui, and Alexandre d'Aspremont. Model selection through sparse maximum likelihood estimation for multivariate G aussian or binary data. J. Mach. Learn. Res., 9: 0 485--516, 2008. ISSN 1532-4435

  3. [3]

    Cand\`es

    Ma gorzata Bogdan, Ewout van den Berg, Chiara Sabatti, Weijie Su, and Emmanuel J. Cand\`es. S LOPE ---adaptive variable selection via convex optimization. Ann. Appl. Stat., 9 0 (3): 0 1103--1140, 2015. ISSN 1932-6157. doi:10.1214/15-AOAS842. URL https://doi.org/10.1214/15-AOAS842

  4. [4]

    Pattern recovery by SLOPE

    Małgorzata Bogdan, Xavier Dupuis, Piotr Graczyk, Bartosz Kołodziejek, Tomasz Skalski, Patrick Tardivel, and Maciej Wilczyński. Pattern recovery by SLOPE . arXiv:2203.12086, 2022

  5. [5]

    Bondell and Brian J

    Howard D. Bondell and Brian J. Reich. Simultaneous regression shrinkage, variable selection, and supervised clustering of predictors with OSCAR . Biometrics, 64 0 (1): 0 115--123, 322--323, 2008. ISSN 0006-341X. doi:10.1111/j.1541-0420.2007.00843.x. URL https://doi.org/10.1111/j.1541-0420.2007.00843.x

  6. [6]

    Parrilo, and Alan S

    Venkat Chandrasekaran, Benjamin Recht, Pablo A. Parrilo, and Alan S. Willsky. The convex geometry of linear inverse problems. Found. Comput. Math., 12 0 (6): 0 805--849, 2012. ISSN 1615-3375. doi:10.1007/s10208-012-9135-7. URL https://doi.org/10.1007/s10208-012-9135-7

  7. [7]

    Patrick Danaher, Pei Wang, and Daniela M. Witten. The joint graphical lasso for inverse covariance estimation across multiple classes. J. R. Stat. Soc. Ser. B. Stat. Methodol., 76 0 (2): 0 373--397, 2014. ISSN 1369-7412. doi:10.1111/rssb.12033. URL https://doi.org/10.1111/rssb.12033

  8. [8]

    Figueiredo and R

    M. Figueiredo and R. Nowak. Ordered W eighted l1 R egularized R egression with S trongly C orrelated C ovariates: T heoretical A spects. In Arthur Gretton and Christian C. Robert, editors, Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, volume 51 of Proceedings of Machine Learning Research, pages 930--938, Cadiz...

Show all 43 references
  1. [9]

    Sparse inverse covariance estimation with the graphical lasso

    Jerome Friedman, Trevor Hastie, and Robert Tibshirani. Sparse inverse covariance estimation with the graphical lasso. Biostatistics, 9: 0 432--441, 2008. doi:10.1093/biostatistics/kxm045. URL https://academic.oup.com/biostatistics/article/9/3/432/224260

  2. [10]

    Estimation of symmetry-constrained G aussian graphical models: application to clustered dense networks

    Xin Gao and H\' e l\`ene Massam. Estimation of symmetry-constrained G aussian graphical models: application to clustered dense networks. J. Comput. Graph. Statist., 24 0 (4): 0 909--929, 2015. ISSN 1061-8600. doi:10.1080/10618600.2014.937811. URL https://doi.org/10.1080/106186...

  3. [11]

    Schneider, T

    P Graczyk, U. Schneider, T. Skalski, and P Tardivel. A U nified F ramework for P attern R ecovery in P enalized and T hresholded E stimation and its G eometry. arXiv:2307.10158, 2023. to appear in J. Optim. Theory Appl

  4. [12]

    Lauritzen

    S ren H jsgaard and Steffen L. Lauritzen. Graphical G aussian models with edge and vertex symmetries. J. R. Stat. Soc. Ser. B Stat. Methodol., 70 0 (5): 0 1005--1027, 2008. ISSN 1369-7412. doi:10.1111/j.1467-9868.2008.00666.x. URL https://doi.org/10.1111/j.1467-9868.2008.00666.x

  5. [13]

    Anti-sparse coding for approximate nearest neighbor search

    Hervé Jégou, Teddy Furon, and Jean-Jacques Fuchs. Anti-sparse coding for approximate nearest neighbor search. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2029--2032, 2012. doi:10.1109/ICASSP.2012.6288307

  6. [14]

    Seyoung Kim and Eric P. Xing. Statistical estimation of correlated genome associations to a quantitative trait network. PLOS Genetics, 5 0 (8): 0 1--18, 08 2009. doi:10.1371/journal.pgen.1000587. URL https://doi.org/10.1371/journal.pgen.1000587

  7. [15]

    Sparsistency and rates of convergence in large covariance matrix estimation

    Clifford Lam and Jianqing Fan. Sparsistency and rates of convergence in large covariance matrix estimation. Ann. Statist., 37 0 (6B): 0 4254--4278, 2009. ISSN 0090-5364. doi:10.1214/09-AOS720. URL https://doi.org/10.1214/09-AOS720

  8. [16]

    Locally associated graphical models and mixed convex exponential families

    Steffen Lauritzen and Piotr Zwiernik. Locally associated graphical models and mixed convex exponential families. Ann. Statist., 50 0 (5): 0 3009--3038, 2022. ISSN 0090-5364. doi:10.1214/22-aos2219. URL https://doi.org/10.1214/22-aos2219

  9. [17]

    Penalized composite likelihood for colored graphical G aussian models

    Qiong Li, Xiaoying Sun, Nanwei Wang, and Xin Gao. Penalized composite likelihood for colored graphical G aussian models. Stat. Anal. Data Min., 14 0 (4): 0 366--378, 2021. ISSN 1932-1864. doi:10.1002/sam.11530. URL https://doi.org/10.1002/sam.11530

  10. [18]

    Learning G aussian graphical models with ordered weighted _1 regularization

    Cody Mazza-Anthony, Bogdan Mazoure, and Mark Coates. Learning G aussian graphical models with ordered weighted _1 regularization. IEEE Trans. Signal Process., 69: 0 489--499, 2021. ISSN 1053-587X. doi:10.1109/TSP.2020.3038480. URL https://doi.org/10.1109/TSP.2020.3038480

  11. [19]

    High-dimensional graphs and variable selection with the lasso

    Nicolai Meinshausen and Peter B \"u hlmann. High-dimensional graphs and variable selection with the lasso. Ann.\ Statist., 2006

  12. [20]

    Renato Negrinho and Andr\' e F. T. Martins. Orbit regularization. In Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc., 2014. URL https://proceedings.neurips.cc...

  13. [21]

    Model selection in gaussian graphical models: High-dimensional consistency of boldmath ell\_1-regularized mle

    Garvesh Raskutti, Bin Yu, Martin J Wainwright, and Pradeep Ravikumar. Model selection in gaussian graphical models: High-dimensional consistency of boldmath ell\_1-regularized mle. In D. Koller, D. Schuurmans, Y. Bengio, and L. Bottou, editors, Advances in Neural Information P...

  14. [22]

    Wainwright, Garvesh Raskutti, and Bin Yu

    Pradeep Ravikumar, Martin J. Wainwright, Garvesh Raskutti, and Bin Yu. High-dimensional covariance estimation by minimizing _1 -penalized log-determinant divergence . Electronic Journal of Statistics, 5 0 (none): 0 935 -- 980, 2011. doi:10.1214/11-EJS631. URL https://doi.org/1...

  15. [23]

    Riccobello, M

    R. Riccobello, M. Bogdan, G. Bonaccolto, P.J. Kremer, S. Paterlini, and P. Sobczyk. Sparse G raphical M odelling via the S orted l_1 - N orm. arXiv:2204.10403, 2022

  16. [24]

    Tyrrell Rockafellar

    R. Tyrrell Rockafellar. Convex analysis. Princeton Mathematical Series, No. 28. Princeton University Press, Princeton, NJ, 1970

  17. [25]

    Tyrrell Rockafellar and Roger J.-B

    R. Tyrrell Rockafellar and Roger J.-B. Wets. Variational analysis, volume 317 of Grundlehren der mathematischen Wissenschaften [Fundamental Principles of Mathematical Sciences]. Springer-Verlag, Berlin, 1998

  18. [26]

    Dependence in elliptical partial correlation graphs

    David Rossell and Piotr Zwiernik. Dependence in elliptical partial correlation graphs. Electron. J. Stat., 15 0 (2): 0 4236--4263, 2021. doi:10.1214/21-ejs1891. URL https://doi.org/10.1214/21-ejs1891

  19. [27]

    Rothman, Peter J

    Adam J. Rothman, Peter J. Bickel, Elizaveta Levina, and Ji Zhu. Sparse permutation invariant covariance estimation. Electron. J. Stat., 2: 0 494--515, 2008. doi:10.1214/08-EJS176. URL https://doi.org/10.1214/08-EJS176

  20. [28]

    The geometry of uniqueness, sparsity and clustering in penalized estimation

    Ulrike Schneider and Patrick Tardivel. The geometry of uniqueness, sparsity and clustering in penalized estimation. J. Mach. Learn. Res., 23: 0 Paper No. [331], 36, 2022. ISSN 1532-4435

  21. [29]

    Covariance estimation for distributions with 2+ moments

    Nikhil Srivastava and Roman Vershynin. Covariance estimation for distributions with 2+ moments. Ann. Probab., 41 0 (5): 0 3081--3111, 2013. ISSN 0091-1798. doi:10.1214/12-AOP760. URL https://doi.org/10.1214/12-AOP760

  22. [30]

    Sparsity and smoothness via the fused lasso

    Robert Tibshirani, Michael Saunders, Saharon Rosset, Ji Zhu, and Keith Knight. Sparsity and smoothness via the fused lasso. J. R. Stat. Soc. Ser. B Stat. Methodol., 67 0 (1): 0 91--108, 2005. ISSN 1369-7412. doi:10.1111/j.1467-9868.2005.00490.x. URL https://doi.org/10.1111/j.1...

  23. [31]

    Tibshirani and Jonathan Taylor

    Ryan J. Tibshirani and Jonathan Taylor. The solution path of the generalized lasso. Ann. Statist., 39 0 (3): 0 1335--1371, 2011. ISSN 0090-5364. doi:10.1214/11-AOS878. URL https://doi.org/10.1214/11-AOS878

  24. [32]

    Turlach, William N

    Berwin A. Turlach, William N. Venables, and Stephen J. Wright. Simultaneous variable selection. Technometrics, 47 0 (3): 0 349--363, 2005. ISSN 0040-1706. doi:10.1198/004017005000000139. URL https://doi.org/10.1198/004017005000000139

  25. [33]

    Model selection with low complexity priors

    Samuel Vaiter, Mohammad Golbabaee, Jalal Fadili, and Gabriel Peyr\' e . Model selection with low complexity priors. Inf. Inference, 4 0 (3): 0 230--287, 2015. ISSN 2049-8764. doi:10.1093/imaiai/iav005. URL https://doi.org/10.1093/imaiai/iav005

  26. [34]

    Model consistency of partly smooth regularizers

    Samuel Vaiter, Gabriel Peyr\' e , and Jalal Fadili. Model consistency of partly smooth regularizers. IEEE Trans. Inform. Theory, 64 0 (3): 0 1725--1737, 2018. ISSN 0018-9448. doi:10.1109/TIT.2017.2713822. URL https://doi.org/10.1109/TIT.2017.2713822

  27. [35]

    How close is the sample covariance matrix to the actual covariance matrix? J

    Roman Vershynin. How close is the sample covariance matrix to the actual covariance matrix? J. Theoret. Probab., 25 0 (3): 0 655--686, 2012 a . ISSN 0894-9840. doi:10.1007/s10959-010-0338-z. URL https://doi.org/10.1007/s10959-010-0338-z

  28. [36]

    Introduction to the non-asymptotic analysis of random matrices

    Roman Vershynin. Introduction to the non-asymptotic analysis of random matrices. In Compressed sensing, pages 210--268. Cambridge Univ. Press, Cambridge, 2012 b

  29. [37]

    Waghmare, Tomas Masak, and Victor M

    Kartik G. Waghmare, Tomas Masak, and Victor M. Panaretos. The F unctional G raphical L asso. arXiv:2306.02347, pages 1--48, 2023. to appear in Ann. Statist

  30. [38]

    Wainwright

    Martin J. Wainwright. Sharp thresholds for high-dimensional and noisy sparsity recovery using _1 -constrained quadratic programming ( L asso). IEEE Trans. Inform. Theory, 55 0 (5): 0 2183--2202, 2009. ISSN 0018-9448. doi:10.1109/TIT.2009.2016018. URL https://doi.org/10.1109/TI...

  31. [39]

    Wainwright

    Martin J. Wainwright. High-Dimensional Statistics: A Non-Asymptotic Viewpoint. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 2019

  32. [40]

    Fused multiple graphical lasso

    Sen Yang, Zhaosong Lu, Xiaotong Shen, Peter Wonka, and Jieping Ye. Fused multiple graphical lasso. SIAM J. Optim., 25 0 (2): 0 916--943, 2015. ISSN 1052-6234. doi:10.1137/130936397. URL https://doi.org/10.1137/130936397

  33. [41]

    Model selection and estimation in the G aussian graphical model

    Ming Yuan and Yi Lin. Model selection and estimation in the G aussian graphical model. Biometrika, 94 0 (1): 0 19--35, 2007. ISSN 0006-3444. doi:10.1093/biomet/asm018. URL https://doi.org/10.1093/biomet/asm018

  34. [42]

    The composite absolute penalties family for grouped and hierarchical variable selection

    Peng Zhao, Guilherme Rocha, and Bin Yu. The composite absolute penalties family for grouped and hierarchical variable selection. Ann. Statist., 37 0 (6A): 0 3468--3497, 2009. ISSN 0090-5364. doi:10.1214/07-AOS584. URL https://doi.org/10.1214/07-AOS584

  35. [43]

    Entropic covariance models

    Piotr Zwiernik. Entropic covariance models. arXiv:2306.03590, pages 1--34, 2023. to appear in Ann. Statist

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.