Pith. sign in

REVIEW 6 cited by

Sparse autoencoders reveal selective remapping of visual concepts during adaptation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.05276 v2 pith:PDCYIKFF submitted 2024-12-06 cs.CV cs.LG

classification cs.CVcs.LG
keywords conceptsadaptationmodelchangedownstreamduringfoundationmechanisms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adapting foundation models for specific purposes has become a standard approach to build machine learning systems for downstream applications. Yet, it is an open question which mechanisms take place during adaptation. Here we develop a new Sparse Autoencoder (SAE) for the CLIP vision transformer, named PatchSAE, to extract interpretable concepts at granular levels (e.g., shape, color, or semantics of an object) and their patch-wise spatial attributions. We explore how these concepts influence the model output in downstream image classification tasks and investigate how recent state-of-the-art prompt-based adaptation techniques change the association of model inputs to these concepts. While activations of concepts slightly change between adapted and non-adapted models, we find that the majority of gains on common adaptation tasks can be explained with the existing concepts already present in the non-adapted foundation model. This work provides a concrete framework to train and use SAEs for Vision Transformers and provides insights into explaining adaptation mechanisms.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail

    cs.LG 2025-12 conditional novelty 6.0 of 10

    Reliable concept presence in transformers is concentrated in the extreme high-activation tail of in-concept tokens; thresholding that tail improves concept detection and localization.

  2. Adversarial Attacks Leverage Interference Between Features in Superposition

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Superposition—packing more features than dimensions—is sufficient to create adversarial vulnerability, and attack directions and transferability are predictable from the resulting feature geometry.

  3. Learning Encoding-Decoding Direction Pairs to Unveil Concepts of Influence in Deep Vision Networks

    cs.CV 2025-09 conditional novelty 6.0 of 10

    An unsupervised method, EDDP, jointly learns encoding-decoding direction pairs for concepts in CNN latent spaces, recovering interpretable and influential concepts without labels, validated on synthetic and real data.

  4. CytoSAE: Interpretable Cell Embeddings for Hematology

    cs.CV 2025-07 conditional novelty 6.0 of 10

    CytoSAE learns sparse, expert-validated morphological concepts from blood-cell images that generalize across datasets and can classify AML subtypes at patient level with F1 0.83.

  5. Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering

    cs.CV 2025-06 unverdicted novelty 6.0 of 10

    VS2 constructs steering vectors from sparse SAE features on unlabeled in-domain activations to improve zero-shot accuracy of CLIP models by 0.93-4.12% on CIFAR-100, CUB-200, and Tiny-ImageNet while remaining forward-p...

  6. Mammo-SAE: Interpreting Breast Cancer Concept Learning with Sparse Autoencoders

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Sparse autoencoders trained on Mammo-CLIP features expose a small set of concept-aligned and confounding latent neurons in breast cancer predictions.

Pith tools