L-SWAG, a zero-shot proxy combining inverse gradient standard deviation with layer-wise activation-pattern counts, plus the LIBRA ensemble rule, outperforms prior proxies on transformer and convolutional search spaces on average.
Abdelfattah, Abhinav Mehrotra, Łukasz Dudziak, and Nicholas D
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
L-SWAG: Layer-Sample Wise Activation with Gradients information for Zero-Shot NAS on Vision Transformers
L-SWAG, a zero-shot proxy combining inverse gradient standard deviation with layer-wise activation-pattern counts, plus the LIBRA ensemble rule, outperforms prior proxies on transformer and convolutional search spaces on average.