In CRATE-family transformers, the attention-like MSSA update with skip connection raises the coding rate it was designed to compress, yet layer-averaged SRR still correlates positively with the generalization gap (tau = 0.445) and slightly improves CIFAR-10/100 accuracy when used as a regularizer.
Transformer interpretability beyond attention visualization
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
In CRATE-family transformers, the attention-like MSSA update with skip connection raises the coding rate it was designed to compress, yet layer-averaged SRR still correlates positively with the generalization gap (tau = 0.445) and slightly improves CIFAR-10/100 accuracy when used as a regularizer.