GIM modifies softmax gradients with a temperature adjustment, layer-norm freeze, and gradient normalization to counter attention self-repair, improving the faithfulness of gradient-based LLM attributions.
XAI for Transformers : Better Explanations through Conservative Propagation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
GIM modifies softmax gradients with a temperature adjustment, layer-norm freeze, and gradient normalization to counter attention self-repair, improving the faithfulness of gradient-based LLM attributions.