Pith. sign in

REVIEW 2 cited by

Structured Sparsity Inducing Adaptive Optimizers for Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.03869 v2 pith:WI3LIXEX submitted 2021-02-07 cs.LG math.OC

classification cs.LGmath.OC
keywords proximaladaptivesparsitygroupsinducingmethodmethodsoperators
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The parameters of a neural network are naturally organized in groups, some of which might not contribute to its overall performance. To prune out unimportant groups of parameters, we can include some non-differentiable penalty to the objective function, and minimize it using proximal gradient methods. In this paper, we derive the weighted proximal operator, which is a necessary component of these proximal methods, of two structured sparsity inducing penalties. Moreover, they can be approximated efficiently with a numerical solver, and despite this approximation, we prove that existing convergence guarantees are preserved when these operators are integrated as part of a generic adaptive proximal method. Finally, we show that this adaptive method, together with the weighted proximal operators derived here, is indeed capable of finding solutions with structure in their sparsity patterns, on representative examples from computer vision and natural language processing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Weight Factorization: Sparse Learning Through the Lens of Artificial Symmetries

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Factorizing weights into D≥2 multiplicative factors and applying L2 weight decay induces a non-convex sparse L2/D penalty, and with tailored initialization and learning rates, achieves superior sparsity-accuracy tradeoffs.

  2. Meta-Sparsity: Learning Optimal Sparse Structures in Multi-task Networks through Meta-learning

    cs.LG 2025-01 reject novelty 4.0 of 10

    Meta-sparsity meta-learns the group-lasso penalty strength lambda via MAML, producing channel-sparse shared backbones for multi-task networks.

Pith tools