Sparsity-guided distillation enables replacing attention layers in ViTs with simpler sequential modules, with sparser layers showing smaller performance drops.
In: First Conference on Language Modeling (2024),https://openreview
3 Pith papers cite this work. Polarity classification is still indexing.
3
Pith papers citing it
years
2026 3representative citing papers
citing papers explorer
-
From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation
Sparsity-guided distillation enables replacing attention layers in ViTs with simpler sequential modules, with sparser layers showing smaller performance drops.
- LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models
- Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents