ImageNet classes in LLaVA-Next are linearly decodable and steerable in the residual stream, and multimodal SAEs reveal that visual and textual features become increasingly shared in deeper layers.
Understanding the role of individual units in a deep neural network
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Line of Sight: On Linear Representations in VLLMs
ImageNet classes in LLaVA-Next are linearly decodable and steerable in the residual stream, and multimodal SAEs reveal that visual and textual features become increasingly shared in deeper layers.