REVIEW 4 cited by
SaiT: Sparse Vision Transformers through Adaptive Token Pruning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
While vision transformers have achieved impressive results, effectively and efficiently accelerating these models can further boost performances. In this work, we propose a dense/sparse training framework to obtain a unified model, enabling weight sharing across various token densities. Thus one model offers a range of accuracy and throughput tradeoffs for different applications. Besides, we introduce adaptive token pruning to optimize the patch token sparsity based on the input image. In addition, we investigate knowledge distillation to enhance token selection capability in early transformer modules. Sparse adaptive image Transformer (SaiT) offers varying levels of model acceleration by merely changing the token sparsity on the fly. Specifically, SaiT reduces the computation complexity (FLOPs) by 39% - 43% and increases the throughput by 67% - 91% with less than 0.5% accuracy loss for various vision transformer models. Meanwhile, the same model also provides the zero accuracy drop option by skipping the sparsification step. SaiT achieves better accuracy and computation tradeoffs than state-of-the-art transformer and convolutional models.
Forward citations
Cited by 4 Pith papers
-
Token Cropr: Faster ViTs for Quite a Few Tasks
Cropr uses removable auxiliary heads to learn task-relevant token pruning in ViTs, achieving 1.5-4x speedups with small accuracy drops across image classification, semantic segmentation, and object detection.
-
Importance-Based Token Merging for Efficient Image and Video Generation
A token-merging method that anchors computation on high-CFG-importance tokens improves generation quality at fixed inference speedups.
-
Fusion: A Framework for Unified Sequential Token AdaptatIon in VisiOn TraNsformers
Staged merge o exit o prune ordering plus lightweight routers lets adaptive ViTs combine three efficiency axes without the accuracy collapse of parallel composition.
-
Back to Fundamentals: Low-Level Visual Features Guided Progressive Token Pruning
LVTP prunes transformer tokens during semantic segmentation using multi-scale Tsallis entropy with Sobel edge guidance, reporting 20-46% FLOP reduction with a few mIoU points lost and no retraining.
Discussion (0). Continue with ORCID to comment.