Keeping 75% of FFN channels idle (linear) lets RePaViT merge two large projections into one, reducing inference FLOPs by ~41% and latency by ~40% on large ViTs.
Token merging: Your vit but faster
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
Keeping 75% of FFN channels idle (linear) lets RePaViT merge two large projections into one, reducing inference FLOPs by ~41% and latency by ~40% on large ViTs.