Real-Time AttentionBender enables live, granular manipulation of self-attention, cross-attention, and feed-forward layers in video diffusion transformers during generation.
AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We present AttentionBender, a tool that manipulates cross-attention in Video Diffusion Transformers to help artists probe the internal mechanics of black-box video generation. While generative outputs are increasingly realistic, prompt-only control limits artists' ability to build intuition for the model's material process or to work beyond its default tendencies. Using an autobiographical research-through-design approach, we built on Network Bending to design AttentionBender, which applies 2D transforms (rotation, scaling, translation, etc.) to cross-attention maps to modulate generation. We assess AttentionBender by visualizing 4,500+ video generations across prompts, operations, and layer targets. Our results suggest that cross-attention is highly entangled: targeted manipulations often resist clean, localized control, producing distributed distortions and glitch aesthetics over linear edits. AttentionBender contributes a tool that functions both as an Explainable AI style probe of transformer attention mechanisms, and as a creative technique for producing novel aesthetics beyond the model's learned representational space.
fields
cs.GR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Real-Time AttentionBender: Granular Interactive Network Bending of Video Diffusion Transformers
Real-Time AttentionBender enables live, granular manipulation of self-attention, cross-attention, and feed-forward layers in video diffusion transformers during generation.