Pith. sign in

Quantized Sparse Weight Decomposition for Neural Network Compression

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

In this paper, we introduce a novel method of neural network weight compression. In our method, we store weight tensors as sparse, quantized matrix factors, whose product is computed on the fly during inference to generate the target model's weights. We use projected gradient descent methods to find quantized and sparse factorization of the weight tensors. We show that this approach can be seen as a unification of weight SVD, vector quantization, and sparse PCA. Combined with end-to-end fine-tuning our method exceeds or is on par with previous state-of-the-art methods in terms of the trade-off between accuracy and model size. Our method is applicable to both moderate compression regimes, unlike vector quantization, and extreme compression regimes.

fields

cs.LG 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Sparse Weight Decomposition for Efficient Circuit Extraction

cs.LG · 2026-08-04 · conditional · novelty 6.0

Sparse Weight Decomposition reparameterizes transformer weight matrices into sparse factors whose bottleneck units support efficient circuit extraction with less data and sparser circuits than learned sparse baselines.

citing papers explorer

Showing 1 of 1 citing paper.

  • Sparse Weight Decomposition for Efficient Circuit Extraction cs.LG · 2026-08-04 · conditional · none · ref 12 · internal anchor

    Sparse Weight Decomposition reparameterizes transformer weight matrices into sparse factors whose bottleneck units support efficient circuit extraction with less data and sparser circuits than learned sparse baselines.