REVIEW 12 cited by
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Transformer networks have revolutionized NLP representation learning since they were introduced. Though a great effort has been made to explain the representation in transformers, it is widely recognized that our understanding is not sufficient. One important reason is that there lack enough visualization tools for detailed analysis. In this paper, we propose to use dictionary learning to open up these "black boxes" as linear superpositions of transformer factors. Through visualization, we demonstrate the hierarchical semantic structures captured by the transformer factors, e.g., word-level polysemy disambiguation, sentence-level pattern formation, and long-range dependency. While some of these patterns confirm the conventional prior linguistic knowledge, the rest are relatively unexpected, which may provide new insights. We hope this visualization tool can bring further knowledge and a better understanding of how transformer networks work. The code is available at https://github.com/zeyuyun1/TransformerVis
Forward citations
Cited by 12 Pith papers
-
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
Attribution-based Parameter Decomposition splits a network's parameters into faithful, minimal, and simple components and recovers ground-truth mechanisms in toy models of superposition and compressed computation.
-
TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series Models
TimeSAE trains a sparse autoencoder with counterfactual and consistency losses to explain black-box time series predictions, claiming better faithfulness and out-of-distribution robustness than eight baselines.
-
Learning Encoding-Decoding Direction Pairs to Unveil Concepts of Influence in Deep Vision Networks
An unsupervised method, EDDP, jointly learns encoding-decoding direction pairs for concepts in CNN latent spaces, recovering interpretable and influential concepts without labels, validated on synthetic and real data.
-
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
A self-supervision method makes multimodal LLMs align their input image embeddings with the model's own refined internal representations, improving visual QA scores over LLaVA baselines.
-
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
Reasoning language models are systematically overconfident, deeper reasoning makes them more overconfident, and a two-stage introspective prompting method improves calibration for some models.
-
Attention-Only Transformers via Unrolled Subspace Denoising
An attention-only transformer, derived as unrolled subspace denoising, provably multiplies token signal-to-noise ratio by a fixed factor per layer and roughly matches GPT-2 and ViT on small benchmarks.
-
LLM Pretraining with Continuous Concepts
A language model trained to predict and interleave teacher-derived SAE concepts into its hidden states beats plain next-token prediction and knowledge distillation on several benchmarks.
-
Obfuscated Activations Bypass LLM Latent-Space Defenses
Obfuscation attacks that jointly optimize for target behavior and for low monitor scores bypass sparse autoencoders, probes, and OOD detectors on LLMs, while performance degrades mainly on hard tasks like writing correct SQL.
-
InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders
Sparse autoencoders trained on ESM-2 recover thousands of interpretable features that align with Swiss-Prot concepts and can influence sequence generation.
-
REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
REFLEX improves explainable fact-checking by using verdict-anchored style control and self-disagreement signals to disentangle fact from style in LLM outputs, achieving SOTA results with minimal self-refined samples.
-
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
A new patch-aligned pretraining loss improves fine-grained vision-language alignment and grounding in multimodal LLMs.
-
A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions
A review that organizes LLM uncertainty quantification into token-level, self-verbalized, semantic-similarity, and mechanistic interpretability categories.
Discussion (0). Continue with ORCID to comment.