EET speeds up vision transformers for fine-grained image retrieval by pruning background tokens and using teacher-student distillation to preserve accuracy, cutting latency by 42.7% with little or no drop in retrieval quality.
Hypergraph-induced semantic tuplet loss for deep metric learning,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Rethinking Vision Transformer for Large-Scale Fine-Grained Image Retrieval
EET speeds up vision transformers for fine-grained image retrieval by pruning background tokens and using teacher-student distillation to preserve accuracy, cutting latency by 42.7% with little or no drop in retrieval quality.