REVIEW 6 cited by
Sparse autoencoders reveal selective remapping of visual concepts during adaptation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Adapting foundation models for specific purposes has become a standard approach to build machine learning systems for downstream applications. Yet, it is an open question which mechanisms take place during adaptation. Here we develop a new Sparse Autoencoder (SAE) for the CLIP vision transformer, named PatchSAE, to extract interpretable concepts at granular levels (e.g., shape, color, or semantics of an object) and their patch-wise spatial attributions. We explore how these concepts influence the model output in downstream image classification tasks and investigate how recent state-of-the-art prompt-based adaptation techniques change the association of model inputs to these concepts. While activations of concepts slightly change between adapted and non-adapted models, we find that the majority of gains on common adaptation tasks can be explained with the existing concepts already present in the non-adapted foundation model. This work provides a concrete framework to train and use SAEs for Vision Transformers and provides insights into explaining adaptation mechanisms.
Forward citations
Cited by 6 Pith papers
-
The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail
Reliable concept presence in transformers is concentrated in the extreme high-activation tail of in-concept tokens; thresholding that tail improves concept detection and localization.
-
Adversarial Attacks Leverage Interference Between Features in Superposition
Superposition—packing more features than dimensions—is sufficient to create adversarial vulnerability, and attack directions and transferability are predictable from the resulting feature geometry.
-
Learning Encoding-Decoding Direction Pairs to Unveil Concepts of Influence in Deep Vision Networks
An unsupervised method, EDDP, jointly learns encoding-decoding direction pairs for concepts in CNN latent spaces, recovering interpretable and influential concepts without labels, validated on synthetic and real data.
-
CytoSAE: Interpretable Cell Embeddings for Hematology
CytoSAE learns sparse, expert-validated morphological concepts from blood-cell images that generalize across datasets and can classify AML subtypes at patient level with F1 0.83.
-
Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
VS2 constructs steering vectors from sparse SAE features on unlabeled in-domain activations to improve zero-shot accuracy of CLIP models by 0.93-4.12% on CIFAR-100, CUB-200, and Tiny-ImageNet while remaining forward-p...
-
Mammo-SAE: Interpreting Breast Cancer Concept Learning with Sparse Autoencoders
Sparse autoencoders trained on Mammo-CLIP features expose a small set of concept-aligned and confounding latent neurons in breast cancer predictions.
Discussion (0). Sign in to comment.