A post-training pipeline clusters MLP activations in pretrained vision transformers and extracts overlapping expert subnetworks, cutting MACs by up to 36% and parameters by up to 32% while retaining about 98% accuracy after fine-tuning.
Cnn mixture- of-depths
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks
A post-training pipeline clusters MLP activations in pretrained vision transformers and extracts overlapping expert subnetworks, cutting MACs by up to 36% and parameters by up to 32% while retaining about 98% accuracy after fine-tuning.