REVIEW 10 cited by
What the DAAM: Interpreting Stable Diffusion Using Cross Attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large-scale diffusion neural networks represent a substantial milestone in text-to-image generation, but they remain poorly understood, lacking interpretability analyses. In this paper, we perform a text-image attribution analysis on Stable Diffusion, a recently open-sourced model. To produce pixel-level attribution maps, we upscale and aggregate cross-attention word-pixel scores in the denoising subnetwork, naming our method DAAM. We evaluate its correctness by testing its semantic segmentation ability on nouns, as well as its generalized attribution quality on all parts of speech, rated by humans. We then apply DAAM to study the role of syntax in the pixel space, characterizing head--dependent heat map interaction patterns for ten common dependency relations. Finally, we study several semantic phenomena using DAAM, with a focus on feature entanglement, where we find that cohyponyms worsen generation quality and descriptive adjectives attend too broadly. To our knowledge, we are the first to interpret large diffusion models from a visuolinguistic perspective, which enables future lines of research. Our code is at https://github.com/castorini/daam.
Forward citations
Cited by 10 Pith papers
-
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
ConceptAttention shows that linear projections in the output space of DiT attention layers yield sharper concept-localizing saliency maps than cross-attention maps, reaching state-of-the-art zero-shot segmentation.
-
SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders
A supervised sparse autoencoder binds each concept to a single neuron, letting Stable Diffusion erase a concept by steering one latent.
-
TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
TARA adds token-focused masking and a token alignment loss to LoRA adapters, allowing several independently trained personalized adapters to be composed with less identity loss and feature leakage.
-
Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback
InnerControl trains lightweight probes on intermediate UNet features to enforce control alignment throughout the denoising trajectory, improving controllability for edges and depth.
-
Diffusion Counterfactual Generation with Semantic Abduction
Diffusion-based causal image counterfactuals with semantic abduction improve identity preservation at a small cost in intervention effectiveness, demonstrated on Morpho-MNIST, CelebA-HQ, and mammogram artifact removal.
-
Spatial Balancing: Designing an LLM-Powered Spatial Externalization Interface for Iterative Science Communication Writing
SpatialBalancing is a system that turns revision trade-offs into spatial navigation so writers can iteratively balance scientific exposition and narrative engagement with LLM assistance.
-
Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts
A method and case study for visualizing cross-attention maps in Wan video diffusion transformers, showing token-region alignment over time and their use as artistic material.
-
TimeMachine: Fine-Grained Facial Age Editing with Identity Preservation
TimeMachine proposes a diffusion model with age-aware cross-attention and a latent age classifier, plus a 1M-image HFFA dataset, claiming SOTA age editing with identity preservation.
-
Unsupervised Class Generation to Expand Semantic Segmentation Datasets
A Stable Diffusion and SAM pipeline generates masked cutouts for novel classes, and mixing them into synthetic source images lets UDA segmentation models learn those classes without modifying the simulator or algorithm.
-
Disentangling Granularity: An Implicit Inductive Bias in Factorized VAEs
A V-shaped pattern in the training objective of factorized VAEs is attributed to a tunable 'disentangling granularity', but the effect is confounded because coarser granularity removes penalty terms by construction.
Discussion (0). Continue with ORCID to comment.