Pith. sign in

REVIEW 2 cited by

Analyzing Redundancy in Pretrained Transformer Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.04010 v2 pith:RQSUB7MA submitted 2020-04-08 cs.CL cs.LG

classification cs.CLcs.LG
keywords redundancymodelsanalysisneuronspretrainedacrossanalyzingapplicability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer-based deep NLP models are trained using hundreds of millions of parameters, limiting their applicability in computationally constrained environments. In this paper, we study the cause of these limitations by defining a notion of Redundancy, which we categorize into two classes: General Redundancy and Task-specific Redundancy. We dissect two popular pretrained models, BERT and XLNet, studying how much redundancy they exhibit at a representation-level and at a more fine-grained neuron-level. Our analysis reveals interesting insights, such as: i) 85% of the neurons across the network are redundant and ii) at least 92% of them can be removed when optimizing towards a downstream task. Based on our analysis, we present an efficient feature-based transfer learning procedure, which maintains 97% performance while using at-most 10% of the original neurons.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sequence Complementor: Complementing Transformers For Time Series Forecasting with Learnable Sequences

    cs.LG 2025-01 conditional novelty 5.0 of 10

    Adding three learnable 'complementor' sequences to each input channel of a PatchTST-style transformer lowers forecasting error by 1-5 percent on standard benchmarks, though the stated information-theoretic guarantee d...

  2. ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation

    cs.LG 2024-11 conditional novelty 5.0 of 10

    A post-training routing method that uses self-distillation to let frozen pretrained Transformers process only a subset of parameters and tokens, cutting active compute by 20 to 50 percent.

Pith tools