CaFE uses attribution-based effective receptive fields to explain sparse autoencoder features in vision transformers, recovering activations better than activation-ranked patches.
Sparse autoencoders reveal selective remapping of visual concepts during adaptation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Causal Interpretation of Sparse Autoencoder Features in Vision
CaFE uses attribution-based effective receptive fields to explain sparse autoencoder features in vision transformers, recovering activations better than activation-ranked patches.