A noise-allocation training trick learns token relevance scores that can prune vision transformer tokens at test time, beating some baselines in some regimes but not all claimed settings.
Computation of channel capacity and rate- distortion functions
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Training Noise Token Pruning
A noise-allocation training trick learns token relevance scores that can prune vision transformer tokens at test time, beating some baselines in some regimes but not all claimed settings.