Pith. sign in

REVIEW 15 cited by

What Do Compressed Deep Neural Networks Forget?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.05248 v3 pith:2OYDZ5OD submitted 2019-11-13 cs.LG cs.AIcs.CVcs.HCstat.ML

classification cs.LGcs.AIcs.CVcs.HCstat.ML
keywords compressiondeepneuralperformancecompresseddatadifferentimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep neural network pruning and quantization techniques have demonstrated it is possible to achieve high levels of compression with surprisingly little degradation to test set accuracy. However, this measure of performance conceals significant differences in how different classes and images are impacted by model compression techniques. We find that models with radically different numbers of weights have comparable top-line performance metrics but diverge considerably in behavior on a narrow subset of the dataset. This small subset of data points, which we term Pruning Identified Exemplars (PIEs) are systematically more impacted by the introduction of sparsity. Compression disproportionately impacts model performance on the underrepresented long-tail of the data distribution. PIEs over-index on atypical or noisy images that are far more challenging for both humans and algorithms to classify. Our work provides intuition into the role of capacity in deep neural networks and the trade-offs incurred by compression. An understanding of this disparate impact is critical given the widespread deployment of compressed models in the wild.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI

    eess.AS 2026-05 accept novelty 7.0 of 10

    The paper delivers a unified framework for fairness in speech technologies by formalizing seven definitions, organizing research into three paradigms, diagnosing pipeline-specific biases, and mapping mitigations to th...

  2. Wake Vision: A Tailored Dataset and Benchmark Suite for TinyML Computer Vision Applications

    cs.CV 2024-05 unverdicted novelty 7.0 of 10

    Wake Vision pipeline produces a 6M-image person detection dataset for TinyML with 2.2% label error, improving model accuracy up to 6.6% over prior VWW benchmark across architectures and subsets.

  3. QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Quantization leaves refusal and multiple-choice bias checks flat while open-ended stereotype endorsement remains high (~24–27% under an independent judge), a gap standard safety evaluations miss.

  4. On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Systematic experiments reveal that activation steering trades fluency for concept control, is less effective on instruction-tuned models, and that prompting/SFT excel at injection but not removal, with textual metrics...

  5. Sigma-Branch: Hierarchical Single-Path Network Reconstruction for Dynamic Inference with Reduced Active Parameters

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Sigma-Branch converts dense networks into single-path hierarchical trees via activation-based spherical k-means clustering and soft-routing fine-tuning, cutting active parameters 58-60% with under 2pp accuracy loss on...

  6. Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels

    cs.LG 2026-05 conditional novelty 6.0 of 10

    3-bit quantization induces new stereotypical biases in 6-21% of previously unbiased BBQ items across three LLMs, undetected by perplexity increases under 3%, with models declining in 'unknown' responses by 17.4%.

  7. SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation

    cs.LG 2023-10 conditional novelty 6.0 of 10

    SalUn uses gradient-based weight saliency to achieve effective machine unlearning of data, classes, or concepts in image classification and generation, narrowing the gap to exact retraining.

  8. DPIFrame: A Dual-Level Parallelism Acceleration Framework for CTR Model Inference

    cs.DC 2026-06 unverdicted novelty 5.0 of 10

    DPIFrame introduces intra- and inter-module parallel architecture, anticipatory multi-table embedding lookup, and breadth-first GPU stream scheduling, reporting 23x embedding latency reduction and up to 5.83x overall ...

  9. Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI

    cs.LG 2026-05 conditional novelty 5.0 of 10

    Activation-aware pruning preserves perplexity but amplifies bias in LLMs, with 47-59% of previously neutral items developing new stereotypical responses at 70% sparsity.

  10. Bias In, Bias Out? Finding Unbiased Subnetworks in Vanilla Models

    cs.LG 2026-03 unverdicted novelty 5.0 of 10

    BISE extracts bias-free subnetworks from conventionally trained models via pruning, enabling debiased operation without retraining or additional data.

  11. The Uneven Impact of Post-Training Quantization in Machine Translation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Across five LLMs and four quantization methods, 4-bit compression mostly preserves translation quality for high-resource languages, while 2-bit compression disproportionately degrades low-resource and Indic languages,...

  12. Uncertainty-Driven Reliability: Selective Prediction and Trustworthy Deployment in Modern Machine Learning

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A training-dynamics abstention method matches deep ensembles at a fraction of the training cost, and a five-term error budget explains why selective classifiers still fall short of the oracle.

  13. When Token Compression Breaks: Structural Pruning vs. Token Reduction for Robust ViT Segmentation under High Compression

    cs.CV 2026-07 unverdicted novelty 4.0 of 10

    Token compression in ViT segmentation degrades sharply at high ratios due to information loss while structural pruning degrades smoothly; a moderate prune-then-merge pipeline improves the trade-off on ADE20K and Citys...

  14. Explaining How Quantization Disparately Skews a Model

    cs.LG 2025-09 conditional novelty 4.0 of 10

    Quantization exacerbates accuracy disparity across groups via a cascade of weight, logit, and probability changes, and a combination of sampling, weighted loss, and mixed-precision training mitigates it.

  15. Compressed Models are NOT Trust-equivalent to Their Large Counterparts

    cs.CL 2025-08 conditional novelty 4.0 of 10

    Compressed BERT models share at most 67% of their top decision features with BERT-base and show different calibration profiles even at similar accuracy, so accuracy parity does not ensure trust-equivalence.

Pith tools