REVIEW 3 cited by
Attention Interpretability Across NLP Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The attention layer in a neural network model provides insights into the model's reasoning behind its prediction, which are usually criticized for being opaque. Recently, seemingly contradictory viewpoints have emerged about the interpretability of attention weights (Jain & Wallace, 2019; Vig & Belinkov, 2019). Amid such confusion arises the need to understand attention mechanism more systematically. In this work, we attempt to fill this gap by giving a comprehensive explanation which justifies both kinds of observations (i.e., when is attention interpretable and when it is not). Through a series of experiments on diverse NLP tasks, we validate our observations and reinforce our claim of interpretability of attention through manual evaluation.
Forward citations
Cited by 3 Pith papers
-
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
DAGCD improves context faithfulness by amplifying context tokens that attention marks as relevant, scaled by token-level uncertainty.
-
A Room to Roam: Reset Prediction Based on Physical Object Placement for Redirected Walking
A Vision Transformer predicts the number of redirected-walking reset events from a top-down occupancy image of a room, achieving R-squared 0.91 in simulation, and powers a real-time furniture-placement interface.
-
New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing
The thesis shows that randomly masking input tokens during fine-tuning makes post-hoc explanations of NLP models consistently faithful under an erasure-based faithfulness metric.
Discussion (0). Continue with ORCID to comment.